Water level identification method and device based on image fusion and electronic equipment

Through the target instance segmentation and depth estimation model combined with multi-image fusion technology, the missegment problem caused by object occlusion in water level recognition is solved, and more accurate water level detection and dynamic change recognition are achieved to adapt to different lighting conditions.

CN120388172APending Publication Date: 2025-07-29CHINA TOWER CO LTD
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202510435105.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-08
Publication Date
2025-07-29

AI Technical Summary

Technical Problem

In the prior art, water level recognition has low target detection accuracy, insufficient depth estimation accuracy, and insufficient multi-image fusion capability, resulting in inaccurate water level recognition results, especially in the case of object occlusion, which makes it difficult to reflect the dynamic changes of water level and adapt to different lighting conditions.

Method used

The target instance segmentation model is used for water segmentation, and the depth map is generated based on the improved target depth estimation model. The water segmentation mask map of the complete blocked water area is output through the multi-image fusion network model. Finally, the water level line of the water mask area is extracted, and the water level detection report is output based on the comparison results of the water level line and the cordon.

Benefits of technology

It improves the accuracy of water level detection, solves the problem of missegment caused by object occlusion, can more accurately identify the dynamic changes of water level and adapt to different lighting conditions, and improves the accuracy and robustness of water level recognition.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120388172A_ABST
    Figure CN120388172A_ABST
Patent Text Reader

Abstract

The invention discloses a water level identification method and device based on image fusion and electronic equipment, and relates to the field of image water level identification or other related fields, and the method comprises the steps: employing a target instance segmentation model to carry out the segmentation processing of a target image in which a water area exists, outputting a global contour segmentation map including the water area and an initial mask map for completing preliminary segmentation of the water area; performing depth estimation on the target image through an improved target depth estimation model to generate a depth map; inputting the initial mask graph, the global contour segmentation graph and the depth graph into a multi-image fusion network model, and outputting a water area segmentation mask graph for complementing the sheltered water area part; extracting a water level line for calibrating the water area mask area in the water area segmentation mask graph; and outputting a water level detection report based on a comparison result of the water level line and the warning line. According to the invention, the technical problem of low water level detection precision caused by object shielding and wrong segmentation of a water area in the prior art is solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of image recognition technology or other related fields. Specifically, it relates to a water level recognition method, device, and electronic device based on image fusion. Background Art

[0002] With the acceleration of the urbanization process, the tasks of urban flood control and drainage have become increasingly prominent. To construct a complete hydrological, meteorological, geological, environmental, and geographic information system for flood monitoring and early warning, remote sensing data plays an important role in flood monitoring and early warning. The key to flood monitoring and early warning is to timely, accurately, and clearly identify, extract, and track targets, and to obtain water level information in real time, accurately, and quickly, and to evaluate flood risks.

[0003] In related technologies, water level warning is a key content of flood warning, and it is necessary to accurately obtain water level information and evaluate flood risks. With the continuous improvement of water conservancy informatization construction, the acquisition of water level information mainly relies on remote sensing images or river video monitoring equipment. In the field of water level monitoring, water level recognition technologies based on image processing, machine learning, and deep learning have gradually become research hotspots. Intelligent image recognition uses an automatic monitoring method to obtain relevant image information of water level monitoring points in real time, which is intuitive and visible, and has the advantage of non-contact automatic detection, and can effectively handle various water level monitoring problems under complex working conditions.

[0004] However, the current water level recognition method based on images has the following technical problems: (1) The target detection accuracy is low, it is difficult to accurately detect large particle targets, such as large target waters, and there is an over-segmentation problem for waters blocked by trees; (2) The accuracy of depth estimation is insufficient, and it is difficult to accurately estimate depth information; (3) The multi-image fusion ability is insufficient, and it is difficult to fuse different image information; (4) The water level recognition result is inaccurate, it is difficult to accurately reflect the dynamic change of the water level, and it has poor adaptability to different lighting conditions and visual sensor data, and low generalization ability.

[0005] For the above problems, no effective solution has been proposed yet. Summary of the Invention

[0006] Embodiments of the present invention provide a water level recognition method, device, and electronic device based on image fusion to at least solve the technical problem in related technologies that due to the existence of object occlusion and over-segmentation of waters, the water level detection accuracy is low.

[0007] To achieve the above object, according to one aspect of the present application, there is provided a water level recognition method based on image fusion, including: using a target instance segmentation model to perform segmentation processing on a target image detecting the presence of water areas, and outputting a global contour segmentation map including the water areas and an initial mask map that has preliminarily segmented the water areas; performing depth estimation on the target image through an improved target depth estimation model to generate a depth map; inputting the initial mask map, the global contour segmentation map, and the depth map into a multi-image fusion network model to output a water area segmentation mask map that complements the occluded water area part; extracting a water level line that calibrates the water area mask region in the water area segmentation mask map; and outputting a water level detection report based on the comparison result between the water level line and the warning line.

[0008] According to another aspect of the embodiments of the present invention, there is also provided a water level recognition device based on image fusion, including: an image segmentation unit for using a target instance segmentation model to perform segmentation processing on a target image detecting the presence of water areas, and outputting a global contour segmentation map including the water areas and an initial mask map that has preliminarily segmented the water areas; a depth estimation unit for performing depth estimation on the target image through an improved target depth estimation model to generate a depth map; an image fusion unit for inputting the initial mask map, the global contour segmentation map, and the depth map into a multi-image fusion network model to output a water area segmentation mask map that complements the occluded water area part; a water level line extraction unit for extracting a water level line that calibrates the water area mask region in the water area segmentation mask map; and a report output unit for outputting a water level detection report based on the comparison result between the water level line and the warning line.

[0009] According to another aspect of the embodiments of the present invention, there is also provided a computer-readable storage medium, where the computer-readable storage medium includes a stored computer program, and wherein, when the computer program runs, it controls the device where the computer-readable storage medium is located to execute the water level recognition method based on image fusion as described in any one of the above.

[0010] According to another aspect of the embodiments of the present invention, there is also provided an electronic device, including one or more processors and a memory, where the memory is used to store one or more programs, and wherein, when the one or more programs are executed by the one or more processors, the one or more processors are caused to implement the water level recognition method based on image fusion as described in any one of the above.

[0011] According to another aspect of the embodiments of the present invention, there is also provided a computer program product, including a computer program, and when the computer program is executed by a processor, it implements the steps of the water level recognition method based on image fusion as described in any one of the above.

[0012] In the present disclosure, a target instance segmentation model is used to segment a target image detected to have a water area, outputting a global contour segmentation map including the water area and an initial mask map that preliminarily segments the water area. An improved target depth estimation model is used to estimate the depth of the target image, generating a depth map. The initial mask map, the global contour segmentation map, and the depth map are input into a multi-image fusion network model to output a water area segmentation mask map that fills in the occluded water area part. The water level line that calibrates the water area mask region in the water area segmentation mask map is extracted, and based on the comparison result between the water level line and the warning line, a water level detection report is output.

[0013] From the above disclosure, the depth of the target image can be estimated by an improved target depth estimation model, the occluded water area part can be filled in by a multi-image fusion network model, and then the water level line that calibrates the water area mask region in the water area segmentation mask map is extracted and compared with the warning line to determine the water level position, solving the problems of mis-segmentation and missed-segmentation existing in the existing water area segmentation algorithms, improving the accuracy of water level position recognition, and thus solving the technical problem in the related art that due to the existence of object occlusion, there is mis-segmentation of the water area, resulting in low water level detection accuracy. BRIEF DESCRIPTION OF THE DRAWINGS

[0014] The drawings described herein are used to provide a further understanding of the present invention, form a part of this application, and the illustrative embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation to the present invention. In the drawings:

[0015] Figure 1 A hardware structure block diagram of a computer terminal (or mobile device) for implementing a water level recognition method based on image fusion is shown;

[0016] Figure 2 It is a flowchart of an optional water level recognition method based on image fusion according to an embodiment of the present invention;

[0017] Figure 3 It is an overall flowchart of an optional water level position recognition based on monocular depth estimation and multi-image fusion technology according to an embodiment of the present invention;

[0018] Figure 4 It is a schematic diagram of an optional instance segmentation using the Mobile SAM model according to an embodiment of the present invention;

[0019] Figure 5 It is a flowchart of an optional training of a depth estimation model according to an embodiment of the present invention;

[0020] Figure 6 It is a network structure diagram of an optional multi-picture fusion module according to an embodiment of the present invention;

[0021] Figure 7 is a flowchart for determining that the water level exceeds the warning line according to an embodiment of the present invention;

[0022] Figure 8 is a flowchart for calculating the relative distance between the water level warning lines according to an embodiment of the present invention;

[0023] Figure 9 is a schematic diagram of an optional water level recognition device based on image fusion according to an embodiment of the present invention;

[0024] Figure 10 is a structural block diagram of an electronic device according to an embodiment of the present application. Detailed implementation manners

[0025] In order to enable those skilled in the art of the present technology to better understand the solution of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0026] It should be noted that the terms "first", "second", etc. in the specification and claims of the present invention and the above-mentioned drawings are used to distinguish similar objects, and do not have to be used to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments of the present invention described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device including a series of steps or units does not have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.

[0027] For the convenience of those skilled in the art to understand the present invention, some terms or nouns involved in the embodiments of the present invention are explained below:

[0028] Instance segmentation combines object detection and semantic segmentation to identify and segment the instances of each object in the image, improving the segmentation accuracy and robustness.

[0029] Monocular depth estimation is used to infer the depth information of the scene from a single two-dimensional image. With the rise of deep learning, convolutional neural networks (CNNs) enable monocular depth estimation to automatically learn the semantic and geometric features in the image, significantly improving the accuracy.

[0030] Multi-image fusion generates higher-quality or feature-specific images by integrating information from multiple images.

[0031] YOLOv8-Seg is used for instance segmentation tasks. By adding an additional branch on top of detection, it generates segmentation masks. This branch typically includes convolutional layers, upsampling operations, and a decoder for generating masks.

[0032] Mobile SAM is an efficient model for instance segmentation. SAM is a general segmentation model that can segment any object in an image given a prompt (such as a point, box, or text). Mobile SAM achieves similar segmentation performance to SAM by using a smaller image encoder (such as tiny VIT) and other structural optimizations, like reducing the number of parameters and accelerating the inference process, while having a faster inference speed and lower hardware requirements.

[0033] deeplabv3++ is a deep learning model for semantic segmentation and the latest version of the deeplab series. The model captures multi-scale features and generates high-quality segmentation masks by introducing the Atrous Spatial Pyramid Pooling (ASPP for short) module and an additional decoder. The ASPP module can pool features at different scales to better understand and interpret details in the image.

[0034] VIT, Vision Transformer, is an image recognition model based on the Transformer architecture. It divides an image into a series of patches and encodes them as a sequence, then uses the self-attention mechanism to capture the global context information in the image. In Mobile SAM, tiny VIT serves as the image encoder, capable of quickly and efficiently extracting image features to support downstream segmentation tasks.

[0035] It should be noted that the water level recognition method and device based on image fusion in this disclosure can be used in the field of image recognition technology for accurate water level position recognition based on monocular depth estimation and multi-image fusion. Also, in the case of accurate water level position recognition based on monocular depth estimation and multi-image fusion, it can be used in any field other than the image recognition technology field. The application field of the water level recognition method and device based on image fusion in this disclosure is not limited.

[0036] It should be noted that the information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) collected in this disclosure are information and data authorized by the user or fully authorized by all parties. Moreover, the processing of relevant data, such as collection, storage, use, processing, transmission, provision, disclosure, and application, complies with the relevant laws, regulations, and standards of the relevant regions, adopts necessary confidentiality measures, does not violate public order and good customs, and provides corresponding operation entrances for users to choose to authorize or refuse.

[0037] The following embodiments of the present invention can be applied to various systems / applications / devices for water level recognition based on image fusion. The present invention solves the problem of incorrect segmentation in the water area blocked by objects through depth estimation and a multi-picture fusion network. At the same time, when judging the distance between the water level line and the warning line, the present invention can calculate the distance between the water level line and the warning line through a fast matching algorithm, which can not only give an alarm when the water level exceeds the warning line, but also give the relative distance when the water level does not exceed the warning line, and can better warn of flood disasters.

[0038] The present invention will be described in detail below in conjunction with each embodiment.

[0039] Embodiment 1

[0040] According to an embodiment of the present invention, an embodiment of a method for water level recognition based on image fusion is provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions. And although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order than here.

[0041] The embodiment of the method for water level recognition based on image fusion provided in the first embodiment of this application can be executed in a mobile terminal, a computer terminal, or a similar computing device. Figure 1 A hardware structure block diagram of a computer terminal (or mobile device) for implementing the method for water level recognition based on image fusion is shown. As Figure 1 shown, the computer terminal 10 (or mobile device) may include one or more ( Figure 1In the figure, the processor 102 (the processor 102 may include, but is not limited to, a processing device such as a microcontroller unit (MCU) or a field programmable gate array (FPGA)) is shown as 102a, 102b, ……, 102n, a memory 104 for storing data, and a transmission device 106 for communication functions. In addition, it may further include: a display, an input / output interface (I / O interface), a universal serial bus (USB) port (which may be included as one of the ports of the BUS bus), a network interface, a power supply, and / or a camera. Those of ordinary skill in the art can understand that Figure 1 The structure shown is only schematic and does not limit the structure of the above electronic device. For example, the computer terminal 10 may further include more or fewer components than Figure 1 shown in the figure, or have a different configuration from Figure 1 shown in the figure.

[0042] It should be noted that the above one or more processors 102 and / or other data processing circuits can generally be referred to as "data processing circuits" herein. The data processing circuit can be embodied in software, hardware, firmware, or any combination thereof, in whole or in part. In addition, the data processing circuit can be a single independent processing module, or be incorporated in whole or in part into any one of other elements in the computer terminal 10 (or mobile device). As involved in the embodiments of the present application, the data processing circuit is a kind of processor control (such as the selection of a variable resistor terminal path connected to an interface).

[0043] The memory 104 can be used to store software programs and modules of application software, such as the program instructions / data storage device corresponding to the water level recognition method based on image fusion in the embodiments of the present application. The processor 102 executes various functional applications and data processing by running the software programs and modules stored in the memory 104, that is, implements the above-mentioned water level recognition method based on image fusion. The memory 104 may include a high-speed random access memory, and may further include a non-volatile memory, such as one or more magnetic storage devices, a flash memory, or other non-volatile solid-state memories. In some instances, the memory 104 may further include a memory remotely set relative to the processor 102, and these remote memories can be connected to the computer terminal 10 through a network. Examples of the above network include, but are not limited to, the Internet, an enterprise intranet, a local area network, a mobile communication network, and their combinations.

[0044] The transmission device 106 is used to receive or send data via a network. Specific examples of the above-mentioned network may include a wireless network provided by the communication provider of the computer terminal 10. In one example, the transmission device 106 includes a network adapter (Network Interface Controller, NIC), which can be connected to other network devices through a base station so as to communicate with the Internet. In one example, the transmission device 106 can be a Radio Frequency (RF) module, which is used to communicate with the Internet wirelessly.

[0045] The display can be, for example, a touch-screen liquid crystal display (Liquid Crystal Display, LCD for short), which enables users to interact with the user interface of the computer terminal 10 (or mobile device).

[0046] The present invention solves the problems of mis-segmentation and missed-segmentation existing in the existing water area segmentation algorithm. The model training only needs a small amount of annotation data for instance segmentation to meet the water area segmentation requirements of medium and high point scenarios, greatly reducing the cost of data annotation.

[0047] Under the above operating environment, the present application provides a Figure 2 water level recognition method based on image fusion as shown. Figure 2 is a flowchart of an optional water level recognition method based on image fusion according to an embodiment of the present invention, as Figure 2 shown, the method includes the following steps:

[0048] Before performing segmentation processing on a target image detected to have a water area using a target instance segmentation model, a target detection model is pre-used to determine whether the target image contains a water area. Optionally, before performing segmentation processing on a target image detected to have a water area using a target instance segmentation model, it further includes: detecting whether there is a water area in the target image using a preset target detection model, where the target detection model is obtained by modifying multiple feature hierarchical structures of a deep learning-based target detection algorithm, and the modification of the multiple feature hierarchical structures includes: deleting targets in the image that are smaller than a preset target size and retaining targets in the image that are larger than the preset target size.

[0049] It should be noted that the object detection model in this embodiment is based on deep learning and has made targeted modifications to multiple feature hierarchical structures. Specifically, post-processing is performed on the feature maps of certain specific levels to optimize the detection results, with particular attention paid to removing smaller objects while retaining larger objects. Among them, the preset object detection model refers to a modified deep learning model, such as YOLOv8, which is adjusted to better detect water areas. The modification of the preset object detection model is mainly reflected in the feature hierarchical structure, especially the detection ability for different-sized objects is adjusted. In deep learning object detection algorithms, such as YOLOv8, the P3, P4, and P5 layers are different levels of the Feature Pyramid Network (FPN), which are used to detect objects of different scales respectively. The P3 layer is usually used to detect smaller objects, while the P5 layer is mainly used to detect larger objects. In this technical feature, the modification includes only retaining the feature levels for detecting larger objects, such as the P4 and P5 layers, while deleting or not using the P3 layer for detecting smaller objects. The purpose of this is to make the model more focused on detecting water areas, which are usually larger, while reducing unnecessary computational effort and improving the detection speed and efficiency.

[0050] In many cases, the area covered by water areas is large, and some small objects or interferences often do not have a substantial impact on water area detection. Therefore, by setting an object size threshold, the model can automatically filter out those objects smaller than the preset size and avoid their interference with the water area detection results. This filtering mechanism can significantly improve the accuracy of the model, especially when dealing with complex natural scenes, because natural scenes usually contain many small interfering objects. Corresponding to filtering out small objects, retaining objects larger than the preset size helps ensure that water areas can be accurately detected. In the high and middle point monitoring scenario, water areas (such as rivers and lakes) often occupy a large area of the image. Therefore, such a modification can make the model more focused on these large objects, thereby improving the accuracy of water level position recognition.

[0051] In practical applications, especially in the high and middle point monitoring scenario, water area objects are often large, and the multi-scale detection structure of the preset object detection model (such as YOLOv8) may introduce unnecessary complexity when dealing with single-scale objects, resulting in waste of computational resources. By adjusting the feature hierarchical structure and only retaining the P4 and P5 layers for detection, the computational effort of the model when processing images can be reduced, the detection speed can be increased, and since water area objects are usually large, this adjustment will not affect the accuracy of water area detection, thus improving the efficiency of the overall algorithm.

[0052] In step S201, the target instance segmentation model is used to segment the target image detected with water areas, and a global contour segmentation map including the water areas and an initial mask map with the water areas preliminarily segmented are output.

[0053] In step S201, if the model detects water areas, instance segmentation processing will be performed. Instance segmentation means accurately segmenting the contours and their respective categories of each object from the image. The target instance segmentation model can be selected arbitrarily. For example, it can be Mobile SAM or FastSAM, etc. In this embodiment, Mobile SAM is used for illustration. The image is segmented twice by this model: one is segmentation without the constraint of the target detection box to output the contour segmentation map of the whole image, including all parts of the water areas and non-water areas; the other is segmentation with the constraint of the target detection box, specifically for segmenting the water area, and the initial mask map of the water area is output. The mask map is a binary map, where the water area part is marked as 1 or 255, and the non-water area part is 0.

[0054] Optionally, the target instance segmentation model is obtained in the following way: an initial vision model based on the Transformer architecture is constructed through multiple recognition modules introducing the spatial attention mechanism; the initial vision model is used to distill a preset image classification model, and the segmentation parameters of the preset image classification model for the entity objects in the image are adjusted to obtain the target instance segmentation model.

[0055] Among them, when obtaining the target instance segmentation model, two key steps of model construction and training are involved, that is, constructing an initial vision model based on the Transformer architecture, and adjusting the segmentation parameters of the model by distillation to finally obtain the optimized target instance segmentation model. Among them, when constructing the initial vision model based on the Transformer architecture, the spatial attention mechanism is introduced. The spatial attention mechanism is a technology that enhances the model's ability to focus on the key areas of the image. It allows the model, when processing the image, to not only focus on the global information but also perform a more detailed analysis of the details at specific positions in the image. In the Transformer-based model, the self-attention mechanism is the core, which can capture the dependencies between different elements in the input sequence. By introducing the spatial attention module, the model's ability in processing position-sensitive tasks (such as instance segmentation) can be enhanced, ensuring that the model can more accurately locate the water area instances.

[0056] Models based on the Transformer architecture typically consist of multiple encoder and decoder modules. Each module performs self-attention operations and may contain a feed-forward neural network (FFN) to further process information. When constructing a target instance segmentation model, designing an architecture with multiple recognition modules (encoders and decoders) can better handle complex information in images. Each recognition module processes the input features through the self-attention mechanism, enabling a more refined understanding of the objects in the image and the relationships between them.

[0057] When obtaining the target instance segmentation model as described above, a preset image classification model is distilled to adjust the segmentation parameters. Here, distillation is a model compression technique where a smaller model (student model) learns to mimic the output of a larger pre-trained model (teacher model), thus acquiring similar performance to the teacher model but with lower computational costs. The preset image classification model serves as the teacher model, while the constructed target instance segmentation model (based on the Transformer architecture and incorporating a spatial attention mechanism) serves as the student model. During the distillation process, the student model learns the output of the teacher model on the image classification task and adjusts its parameters so that it can better understand the entity objects in the image while maintaining or enhancing its performance on the instance segmentation task.

[0058] Using the self-attention mechanism to process image features and how to more accurately locate the target during the segmentation process. The goal of adjusting the segmentation parameters is to maintain or enhance the segmentation accuracy and robustness while reducing the model complexity. This optimization process usually involves supervising the model output to ensure that the student model not only approximates the teacher model in classification but also performs well in the segmentation task.

[0059] Optionally, the steps of using the target instance segmentation model to segment a target image detected to have water areas, and outputting a global contour segmentation map including the water areas and an initial mask map for the preliminary segmentation of the water areas are as follows: using the target instance segmentation model to perform global segmentation on the target image to obtain a global contour segmentation map including the water areas; using the water area target detection box as a prompt feature, and using the target instance segmentation model to perform water area segmentation on the target image to obtain an initial mask map for the preliminary segmentation of the water areas.

[0060] It should be noted that when using the target instance segmentation model to perform global segmentation on the target image, the target instance segmentation model processes the target image, not only identifying the water area, but also being able to identify and segment other objects and the background in the image, generating a global contour segmentation map. For example, in Mobile SAM, the image is encoded by the Tiny VIT encoder to extract the feature representation of the image. Then, these features are passed to the prompt-based mask decoder to generate segmentation masks corresponding to all instances, including the water area and non-water area objects. The global contour segmentation map provides the segmentation results of all objects in the image, not limited to the water area, which helps with subsequent image fusion and the utilization of depth information.

[0061] After that, during the process of obtaining the initial mask map through preliminary water area segmentation, using the water area target detection box as the prompt feature, the target instance segmentation model is used to perform water area segmentation on the target image: After successful water area target detection, the target detection box (output by models such as YOLOv8, for example) is passed as the prompt feature to the target instance segmentation model. The mask decoder of Mobile SAM can receive these prompt features and focus on generating the segmentation results of the water area. This embodiment utilizes the precise position information provided by the target detection box, making the segmentation process more efficient and accurate. The generated initial mask map is only for the water area. Compared with unconstrained global segmentation, segmentation using the target detection box constraint can avoid misidentifying non-water areas as water areas, improving the accuracy of water area segmentation.

[0062] In the target instance segmentation task, prompt features (Prompts) are particularly important. They can be target boxes, points, lines, or text descriptions, used to guide the model to focus on specific regions or targets in the image. In water area segmentation, using the target detection box as the prompt feature can significantly improve the positioning accuracy of the segmentation model for the water area, avoiding segmenting irrelevant regions, and improving the efficiency and accuracy of the algorithm. In addition, the use of the target detection box reduces the requirement for the model's generalization ability, enabling the model to focus more on learning water area features during training rather than other background features, which has a positive effect on reducing the amount of data required for training and improving the segmentation effect.

[0063] The strategy of combining the target detection and the target instance segmentation model is a manifestation of making full use of their respective advantages. The target detection model can quickly provide the position information of the water area target, while the target instance segmentation model can generate a high-precision water area segmentation mask based on this position information, not only improving the segmentation efficiency, but also reducing the dependence of instance segmentation on a large amount of labeled data. At the same time, using the constraint of the target detection box avoids possible missed segmentation and missegmentation problems in instance segmentation.

[0064] Step S202: Use the improved target depth estimation model to estimate the depth of the target image and generate a depth map.

[0065] In step S202, the improved depth estimation model is used to estimate the depth information of each pixel in the target image. A depth map is an image where the value of each pixel corresponds to the depth of that point from the camera. Adopting the idea of model distillation, it is jointly trained using a large amount of unlabeled water area data from the business scenario and a small amount of labeled open-source depth estimation data (such as NYUv2 and KITTI). First, fine-tune the image encoder, and then fine-tune the depth estimation model to generate a depth map. The depth map is an important basis for solving the problem of tree occlusion of water areas because different depth values can be used to infer the actual distance between the occluding object and the water surface, thereby identifying the occluded water area part.

[0066] Optionally, the target depth estimation model is obtained in the following way: Use the unlabeled water area dataset from the business scenario to adjust the target image encoder; freeze the target image encoder and use the labeled open-source depth estimation dataset to adjust the initial depth estimation model; use the adjusted initial depth estimation model to infer the unlabeled water area data from the business scenario to obtain a depth estimation map; adopt a preset data transformation strategy to transform the obtained depth estimation map to obtain a labeled water area dataset; use the labeled water area dataset to train the specified image encoder in the adjusted preset image classification model to obtain the target depth estimation model.

[0067] The construction process of the target depth estimation model is to utilize the unlabeled data from the business scenario and the labeled open-source depth estimation data to train a high-performance depth estimation model suitable for a specific scenario. First, use a large amount of unlabeled water area image datasets to adjust the target image encoder (such as based on DINOv2). This process can be achieved through semi-supervised learning or self-supervised learning, that is, the model learns the feature representation of the image without depth labels. The main goal of adjusting the image encoder is to enable it to better capture the features of the water area scene and provide high-quality image features for subsequent depth estimation. After the image encoder adjustment is completed, freeze it, that is, no longer update its weights. This means that in the subsequent training, the parameters of the image encoder will remain unchanged, and only other parts of the depth estimation model, such as the depth decoder, will be trained. This can ensure that the features learned by the encoder in the business scenario can be fully utilized, and at the same time, it also avoids overfitting the encoder during the training of the depth estimation model.

[0068] Next, use the labeled open-source depth estimation dataset to adjust an initial depth estimation model. The initial depth estimation model can be a pre-trained model with certain depth estimation capabilities. The adjustment process includes fine-tuning the depth decoder of the model to make it perform better on the labeled depth maps. By training on the open-source dataset, the model can learn general depth estimation rules, which lays a foundation for subsequent applications on self-business scenario data. Then, use the adjusted initial depth estimation model to perform inference on the unlabeled water area dataset in the self-business scenario to generate depth estimation maps. Although these depth estimation maps do not have real depth labels, they already contain certain depth information. To convert these unlabeled depth estimation maps into labeled data available for training, a preset data transformation strategy needs to be adopted. This strategy can include methods such as enhancing the contrast of depth information, normalizing the depth maps, or generating pseudo-depth labels for each pixel point. The purpose is to make the generated depth estimation maps simulate the distribution of real depth maps, so as to be used for training the depth estimation model.

[0069] Finally, through the data transformation strategy, the inferred depth estimation maps are converted into a labeled water area dataset, which means that each image now has a corresponding depth map. Use these datasets to train the specified image encoder in the preset image classification model. It should be noted here that the above-specified image encoder can refer to the encoder in the fine-tuned preset classification image model. For example, the Tiny VIT encoder of the fine-tuned Mobile SAM. Through training, the encoder can learn deeper feature representations on images containing depth information, while the decoder learns how to recover accurate depth maps from these features, thus obtaining the final target depth estimation model.

[0070] Through the training and optimization of the above steps, the target depth estimation model can not only utilize the unlabeled data in the self-business scenario but also combine the open-source depth estimation dataset to generate accurate depth maps, which are used to solve problems such as object occlusion and improve the accuracy of water level position recognition.

[0071] Step S203: Input the initial mask map, the global contour segmentation map, and the depth map into the multi-image fusion network model, and output a water area segmentation mask map that complements the occluded water area part.

[0072] Optionally, after inputting the initial mask map, the global contour segmentation map, and the depth map into the multi-image fusion network model, it further includes: the multi-image fusion network model performs channel splicing processing on the initial mask map, the global contour segmentation map, and the depth map; the multi-image fusion network model performs multiple convolution processes on the spliced image to output a water area segmentation mask map for complementing the occluded water area part. Among them, during the convolution process, a channel attention mechanism and a spatial attention mechanism are introduced. The channel attention mechanism is used to extract the global information of each channel of the image through global average pooling processing and global maximum pooling processing, and then generate the weights of the channels through multiple activation functions. The weights are used to weight each channel of the input feature map to enhance the features of the target channel; the spatial attention mechanism is used to perform global maximum pooling processing and global average pooling processing on the spliced image in the channel dimension to generate a spatial attention map. The spatial attention map is used to identify the spatial positions in the spliced image and enhance the spatial positions; among them, during the process of the multi-image fusion network model performing convolution processing on the image, upsampling processing and downsampling processing are performed on the image.

[0073] This embodiment needs to generate a water area segmentation mask map for complementing the occluded water area part. First, the initial mask map, the global contour segmentation map, and the depth map are spliced in the channel dimension. Here, channel splicing means connecting multiple images or feature maps in the channel dimension to form a multi-channel image. The spliced image has multi-modal information, that is, the preliminary water area segmentation information provided by the initial mask map, the object boundary information given by the global contour segmentation map, and the depth information of the depth map.

[0074] It should be noted that in the convolution process, the multi-image fusion network model enhances the features of the target channels by introducing a channel attention mechanism. The channel attention mechanism extracts the global information of each channel from the entire feature map through global average pooling and global maximum pooling operations, and then generates the weights of this channel through a series of activation functions. These weights are used to weight each channel of the input feature map, enabling the model to automatically identify which channel features are most important for the task and enhancing the contributions of these channels, thereby improving the accuracy of the model for water area segmentation. In the context of image fusion, the channel attention mechanism can highlight the features that are most helpful for water area segmentation in the stitched image, such as the boundary features between the water area and the parts blocked by trees and the channels corresponding to depth information. At the same time, this embodiment also introduces a spatial attention mechanism to identify and enhance specific spatial positions in the stitched image. It performs global maximum pooling and global average pooling on the stitched image in the channel dimension to generate a spatial attention map, which can indicate which positions are crucial for the image fusion task. When dealing with occlusion problems, the spatial attention mechanism can help the model focus on the occluded water area, and by enhancing the feature representation of these areas, improve the model's ability to identify and complete the occluded water area part.

[0075] During the convolution process, downsampling can also be used to reduce the resolution of the feature map and extract higher-level feature representations; upsampling is used to restore the resolution and generate an output of the same size as the input image. These operations are usually combined with skip connections to reintroduce the detailed information lost during the downsampling process into the upsampling process to generate a clearer segmentation mask map. In this embodiment, the purpose of the upsampling and downsampling processes is to generate a high-resolution water area segmentation mask map for completing the occluded water area part to ensure the details and accuracy of the segmentation result.

[0076] Optionally, the multi-image fusion network model is obtained in the following way: a synthetic occlusion dataset and a real occlusion dataset are constructed. The synthetic occlusion dataset refers to pasting a mask of a water area carrying a specified feature object on the edge of the unoccluded water area, providing labels for the unoccluded data, and the specified feature objects include at least: trees, rocks, and moss; the pre-constructed convolutional neural network architecture based on image segmentation is trained using the synthetic occlusion dataset and the real occlusion dataset to obtain the multi-image fusion network model.

[0077] Among them, when constructing the synthetic occlusion dataset, the mask of an object carrying a specific feature (such as a tree, a rock, moss, etc.) is pasted on the originally unoccluded water area image to simulate the occlusion situation that may occur in the actual monitoring scenario. Specifically, the mask pasting process involves selecting regions containing the specified feature objects from other images and placing the contours of these objects on the edge of the unoccluded water area image through image processing techniques (such as image stitching, Alpha blending, etc.) to simulate the occlusion effect.

[0078] In addition, it should be noted that different from the synthetic occlusion dataset, the real occlusion dataset is obtained from the actual monitoring scenario, which naturally contains the occlusion phenomenon. These datasets do not require artificial addition of occlusion, but directly collect images containing occlusion situations, such as water area images occluded by trees, providing more complex and diverse occlusion instances, which helps the model learn the processing strategies in the natural occlusion environment and improve its generalization ability in actual applications.

[0079] Before using the synthetic occlusion dataset and the real occlusion dataset for training, it is necessary to construct a convolutional neural network architecture of a multi-image fusion network, which includes multiple convolutional layers, upsampling layers, and attention mechanism modules (such as spatial attention and channel attention), aiming to process and fuse different types of image inputs, such as depth maps, water area segmentation maps, and full-image contour segmentation maps. During the training process, it can be achieved by providing input images (including depth maps, water area segmentation maps, and full-image contour segmentation maps) and corresponding output labels (i.e., the water area segmentation mask map after filling in the occlusion). The model will learn how to integrate the features in the input images to generate an output image that accurately reflects the water area boundary and fills in the occluded area. During the training process, the use of the synthetic occlusion dataset and the real occlusion dataset is complementary. The former provides a large number of controlled occlusion instances, which helps the model learn basic occlusion recovery skills; the latter provides natural occlusion scenarios, which helps the model improve its performance in the actual complex environment.

[0080] Furthermore, the multi-image fusion network model in this embodiment can also be adaptively replaced with a GAN network.

[0081] Step S204, extract the water level line that calibrates the water area mask region in the water area segmentation mask map.

[0082] Extract the water level line from the water area segmentation mask map that has filled in the occluded part. The water level line is the boundary between the water area mask region and non-water area mask regions such as land, usually in a linear shape. Through computer vision techniques, such as edge detection, contour tracing, etc., the water level line is accurately located and marked on the segmentation mask map.

[0083] Step S205, based on the comparison result between the water level line and the warning line, output a water level detection report.

[0084] In step S205, the water level line is compared with the manually demarcated warning line to determine whether the water level exceeds the warning line. This comparison process can use various methods, such as intersection judgment, distance calculation, etc. If the water level line reaches or exceeds the warning line, the system will immediately generate a water level detection report and send an alarm signal; if the water level line does not exceed the warning line, the relative distance between the water level line and the warning line will be calculated to monitor the trend of water level changes and give early warnings of possible flood disasters. The water level detection report includes information such as the water level status (whether it exceeds the warning line) and the relative position between the water level line and the warning line, which is of great value for water level monitoring and disaster prevention.

[0085] Optionally, the step of outputting the water level detection report based on the comparison result between the water level line and the warning line includes: inputting the starting point of the warning line and the water area segmentation mask map; sampling multiple feature points at equal intervals from the warning line and counting the pixel values corresponding to the multiple feature points in the water area segmentation mask map; in the case where the number of pixel values indicating that the feature points are located within the water area exceeds a preset quantity threshold, it is confirmed that the water level in the water area exceeds the warning line, and the water level line and the warning information are entered into the water level detection report; in the case where the number of pixel values indicating that the feature points are located within the water area does not exceed the preset quantity threshold, it is confirmed that the water level in the water area does not exceed the warning line, and the water level line is entered into the water level detection report.

[0086] Among them, the starting point of the warning line is preset to indicate the position of the warning line in the image coordinate system. The warning line can be represented as a straight line or a curve in the image. In order to accurately evaluate the relationship between the warning line and the water level line, it is necessary to sample multiple feature points at equal intervals from the warning line. These feature points can be evenly distributed along the length of the warning line. For example, 20 points are sampled to comprehensively analyze the contact situation between the warning line and the water area. The selection of the sampling strategy should consider the length of the warning line and the possible changes in the water area to ensure that the sampling points can cover the key areas of the warning line. The sampled feature points need to find the corresponding positions in the water area segmentation mask map and read the pixel values at these positions. The pixel values reflect whether the feature points are located within the water area. For example, a pixel value of 255 indicates that the feature point is within the water area, while a pixel value of 0 indicates that the feature point is outside the water area. The water level detection report is a document generated based on the comparison result, which records the water level status of the water area and the relevant warning information. If the water level exceeds the warning line, the report will include the specific position information of the water level line and the alarm signal so that relevant departments or personnel can immediately take countermeasures. If the water level does not exceed the warning line, the report will only include the position information of the water level line and a possible analysis of the water level trend, providing data support for preventing possible future water level increases.

[0087] It should be further noted that the equal-proportion sampling mentioned in this embodiment refers to evenly distributing sampling points on the warning line. This strategy can ensure that every part of the warning line is taken into account, avoiding detection errors that may be caused by uneven distribution of sampling points. The number and distribution strategy of sampling points should be optimized according to the length and shape of the warning line, as well as the possible variation range of the water area, to ensure the accuracy of the detection results.

[0088] Furthermore, the water level detection report in this embodiment not only includes the position information of the water level line, but also can include data such as water quality status, flow rate, water level change trend, etc., as well as a high-resolution water area segmentation mask map and a schematic diagram of the warning line. The format of the report should be clear and easy to read to ensure that relevant departments can quickly understand the water level status and make timely responses according to the situation. The report should also include the detection date and time, detection equipment information, and possible alarm information for easy recording and tracing. When it is confirmed in the water level detection report that the water level exceeds the warning line, the system should immediately activate the alarm mechanism and send alarm information to relevant departments or personnel, including the specific position of the water level line, the degree of over-warning, and possible subsequent impacts. The alarm mechanism can be in the form of text messages, emails, mobile application notifications, or automated reporting systems to ensure that the alarm information can be conveyed quickly and accurately.

[0089] Through the above steps, the target instance segmentation model can be used to segment the target image with detected water areas, output a global contour segmentation map including the water area and an initial mask map for the preliminary segmentation of the water area, perform depth estimation on the target image through the improved target depth estimation model to generate a depth map, input the initial mask map, global contour segmentation map, and depth map into the multi-image fusion network model, output a water area segmentation mask map that complements the occluded water area part, extract the water level line that calibrates the water area mask region in the water area segmentation mask map, and based on the comparison result between the water level line and the warning line, output a water level detection report. In this embodiment, the improved target depth estimation model can be used to perform depth estimation on the target image, the multi-image fusion network model can be used to complement the occluded water area part, and then by extracting the water level line that calibrates the water area mask region in the water area segmentation mask map and comparing it with the warning line to determine the water level position, the problems of mis-segmentation and missed segmentation existing in the existing water area segmentation algorithms are solved, the accuracy of water level position recognition is improved, and thus the technical problem that the water level detection accuracy is low due to object occlusion and mis-segmentation of the water area in the related technology is solved.

[0090] Optionally, after outputting the water level detection report based on the comparison result between the water level line and the warning line, it further includes: performing downsampling processing on the water area segmentation mask map and reducing the starting point coordinates of the warning line; detecting the direction of the warning line in the water area and performing feature sampling on the reduced warning line; if the warning line is above and below the water area, selecting the intersection points of the sampling points of the warning line and the water area in the Y direction; if the warning line is on the left and right of the water area, selecting the intersection points of the sampling points of the warning line and the water area in the X direction; selecting a predetermined number of intersection points from all the selected intersection points and using the selected intersection points as the initial points, where the selected initial points are the intersection points with the smallest distance from the corresponding sampling points; for each initial point, calculating the distance values between the initial point and each sampling point on the warning line, and determining the minimum distance value between the initial point and all sampling points on the warning line through a binary search strategy; comparing the minimum distance values between the initial points and the sampling points on the warning line, and selecting the minimum distance value as the target warning distance value; if the warning line is above and below the water area, dividing the target warning distance value by the height value of the water area segmentation mask map to obtain the relative distance; or, if the warning line is on the left and right of the water area, dividing the target warning distance value by the width value of the water area segmentation mask map to obtain the relative distance, where the relative distance is used to record the change information of the water level position.

[0091] In this embodiment, the size of the segmentation mask map is reduced through downsampling processing, and at the same time, the starting point coordinates of the warning line are correspondingly reduced, which can significantly reduce the computational complexity of subsequent operations and improve the processing speed. The downsampling ratio needs to be determined according to the accuracy requirements of the warning line and the availability of computing resources, and usually remains within a range that can accurately reflect the relative positions of the water level and the warning line. After the reduction processing, the system needs to accurately determine the direction of the warning line in the water area segmentation map, that is, to determine whether the warning line is located above, below, to the left, or to the right of the water area. For the warning line above and below the water area, the system selects the intersection points of the sampling points of the warning line and the water area in the Y direction; if the warning line is located on the left and right of the water area, the intersection points of the sampling points of the warning line and the water area in the X direction are selected. These intersection points are regarded as the contact points between the warning line and the water area and can accurately reflect the distance relationship between the water level line and the warning line. Next, a predetermined number of intersection points are selected from all the intersection points as the initial points for distance calculation. The selection criterion is that these initial points should be the intersection points with the smallest distance from the sampling points of the warning line to ensure the accuracy of the calculation.

[0092] After obtaining the minimum distance values between all initial points and the warning line sampling points, the system will compare these distance values and select the minimum distance value as the target warning distance value. This distance value can reflect the actual distance between the water level line and the warning line. Finally, divide the target warning distance value by the height value of the water area segmentation mask map (when the warning line is above and below the water area) or the width value (when the warning line is on the left and right of the water area), and the obtained ratio is the relative distance. The relative distance is a standardized index that is not affected by the image resolution and can more intuitively reflect the positional relationship between the water level and the warning line and the change trend of the water level.

[0093] The calculated relative distance is a standardized index that is not affected by the image size and can provide a stable reference value for recording and analyzing the change trend of the water level in the water area. This is very useful in long-term monitoring. By recording the relative distances at different time points, a water level change curve can be plotted, providing data support for predicting the water level trend, assessing flood risks, and water resource management.

[0094] This embodiment can be applied to various water area monitoring scenarios, especially suitable for medium and high elevation monitoring environments, such as remote monitoring points of rivers, lakes, etc. These monitoring points may face occlusion of the water area by natural objects such as trees and rocks. In such scenarios, accurately monitoring whether the water level line exceeds the warning line and calculating the relative distance between the water level line and the warning line is crucial for preventing and managing flood disasters.

[0095] By performing downsampling processing on the water area segmentation mask map and reducing the starting point coordinates of the warning line, this technical feature can ensure the accuracy of comparing the warning line with the water level line while maintaining the calculation efficiency. By accurately detecting the direction of the warning line in the water area and performing feature sampling on the reduced warning line, the relationship between the warning line and the water level line can be analyzed more carefully, avoiding the rough judgment and false alarms that may exist in conventional methods.

[0096] When the water level does not exceed the warning line, the technical feature can calculate the relative distance between the warning line and the water level line. This distance is based on the intersection points of the warning line and the water area segmentation mask map. By selecting the intersection point with the minimum distance from the sampling points as the initial point, and then calculating the minimum distance values between the initial point and each sampling point on the warning line. Quickly determining the minimum distance value through the binary search strategy not only improves the calculation efficiency but also ensures the accuracy of distance measurement. The finally obtained relative distance can provide quantitative information on the change of the water level position, which is of great value for predicting the water level trend and formulating flood control measures.

[0097] The following is a detailed description in combination with another optional specific implementation manner.

[0098] Figure 3It is an overall flowchart of water level position recognition based on monocular depth estimation and multi-image fusion technology according to an embodiment of the present invention. As Figure 3 shown, for the pictures in the monitoring scenario, first, through object detection, it is judged whether there is a water area (such as rivers, lakes, etc.) in the picture. If there is a water area, the original picture is input into Mobile SAM twice. Without any constraints for the first time, the segmentation results of all contours can be obtained. With the water area detection box constraint for the second time, the masked area of the water area segmentation is output. At the same time, monocular depth estimation is performed on the picture to obtain a depth map. Then, the global contour segmentation map, the water area segmentation masked map, and the depth Figure 1 are input into the multi-image fusion generation network. The water area segmentation masked result output by the fusion generation network will automatically complete the water area part blocked by the woods. Finally, first, the complete water level line is obtained from the water area masked area, and then the warning line is compared with the water level line. If the water level line has exceeded the warning line, an alarm will be immediately sent to the platform. If it has not exceeded the warning line, the relative distance between the water level line and the warning line will be calculated and output. The core steps of the embodiment of the present invention include object detection, Mobile SAM segmentation, depth estimation, multi-image fusion module, water level exceeding warning line determination module, and relative distance calculation between the water level line and the warning line. The present invention will be described in detail below in combination with these steps respectively.

[0099] The first step, object detection: The embodiment of the present invention is based on yolov8s to detect the water area in the picture. About 2000 pictures are used to train the model. The size of the input picture during training can be 960*960. Since most of the water area detection boxes in the picture are single large targets and there are almost no dense small targets, the embodiment of the present invention modifies the feature pyramid structure of yolov8. In yolov8, the P3, P4, and P5 layers are different levels of the feature pyramid (Feature Pyramid). The P3 layer is usually used to detect smaller targets, the P4 layer is used to detect medium-sized targets, and the P5 layer is mainly used to detect larger targets. The embodiment of the present invention no longer uses the P3 layer for prediction, which can not only reduce the calculation amount but also enable the model to pay more attention to the P4 and P5 layers and improve the detection accuracy of large targets.

[0100] The second step, using the Mobile SAM model for instance segmentation: Mobile SAM is a lightweight version of SAM, Figure 4 which is a schematic diagram of using the Mobile SAM model for instance segmentation according to an embodiment of the present invention. As Figure 4As shown, compared with the original SAM, the image encoder is replaced by tiny VIT, and other structures are the same as those of SAM. During training, the large model of SAM with VIT-H is used to distill the small model of Mobile SAM to improve its segmentation performance. The speed of Mobile SAM can meet the real-time requirements. To further improve the water area segmentation effect of Mobile SAM in medium and high-point monitoring scenarios, the embodiment of the present invention fine-tunes Mobile SAM. 200 labeled water area segmentation data are used for fine-tuning, and only 2 rounds of iteration are required. Mobile SAM can achieve an ideal water area segmentation effect at medium and high points. In the embodiment of the present invention, Mobile SAM is used for inference twice, and there is only one image encoding in the two inferences. In the prompt-based mask decoding stage, no prompt information is provided in the first inference, and the model will output the result of full-image segmentation. In the second inference, the water area target detection box is provided as prompt information, and the model will output the result of water area segmentation.

[0101] The third step, monocular depth estimation: For the problem of occlusion, it can be judged according to the depth map. If a tree occludes the water area, there will be a large difference in the depth information of the occluded area and the depth information of the water area. The embodiment of the present invention can perform monocular depth estimation based on an improved algorithm of Depth Anything. Depth Anything is a semi-supervised depth estimation algorithm that can be jointly trained using labeled data and unlabeled data.

[0102] Figure 5 It is a flowchart of an optional depth estimation model training according to an embodiment of the present invention. As Figure 5 shown, in the first stage, the image encoder DINOv2 is fine-tuned using a large amount of unlabeled water area data (more than 10W) from the business scenario. In the second stage, the encoder DINOv2 is frozen, and the large depth estimation model is fine-tuned using labeled open-source depth estimation data (such as NYUv2 and KITTI). In the third stage, the fine-tuned deep learning large model is used to infer a large amount of unlabeled water area data from the business scenario, and data transformation is performed on the obtained depth estimation map to obtain a large amount of pseudo-GT data (i.e., labeled data set). In the fourth stage, the small model is fine-tuned. The image encoder of the small model selects the Tiny VIT encoder of Mobile SAM fine-tuned in step 3, and the training data uses the pseudo-GT data obtained in the third stage. The obtained depth estimation small model is the final required model.

[0103] Step 4: Multi-image fusion module. The water area segmentation map, full segmentation map, and depth map are obtained from the previous steps. Since the full segmentation map can provide the contour information of each module in the image, and the depth map can provide the depth information of each module, theoretically, fusing the three-way images can generate an unoccluded water area segmentation map. To reduce the computational load, in this embodiment, the three-way images are all resized to a size of 512, and a multi-image fusion network is designed with reference to the structure of UNET. First, the images are downsampled and then upsampled.

[0104] Figure 6 It is an optional network structure diagram of the multi-image fusion module according to an embodiment of the present invention. As Figure 6 shown, first, the water area segmentation mask map, full image segmentation map, and depth estimation map are concatenated in channels, and then the feature alignment of the three types of images is performed. To successfully align the features of the three types of images, the present invention embodiment adds a spatial attention mechanism and a channel attention mechanism. The channel attention mechanism extracts the global information of each channel through global average pooling and global maximum pooling, and then generates the weights of the channels through a series of activation functions (such as Figure 6 the content indicated by Conv2D*+ReLU in). These weights are used to weight each channel of the input feature map, thereby enhancing the features of important channels. The spatial attention mechanism generates a spatial attention map by performing global maximum pooling and global average pooling on the feature map in the channel dimension. This attention map can identify important spatial positions in the input image, enhance these positions, and thus better capture key spatial information. Combining the channel attention mechanism and the spatial attention mechanism, the network can more accurately focus on the features and regions most important for a specific task, improving the overall performance of the model. The model is mainly trained with synthetic occlusion data (for example, 3000 pieces), supplemented by real occlusion data (for example, 200 pieces). The data synthesis scheme mainly pastes the mask of the tree on the edge of the unoccluded water area, and the unoccluded data provides the gt. The three-way input images are first concatenated in channels, and then the final mask output is obtained through modules such as convolution, upsampling layer, attention mechanism, and upsampling.

[0105] It should be noted that, as Figure 6 shown, during the convolution process of the concatenated image, downsampling processing will be performed ( Figure 6 indicated by MaxPooling*(1 / 2) in), and then after passing through the channel attention module (channel attention mechanism) and the spatial attention module (spatial attention mechanism), upsampling processing is performed on the image ( Figure 6 indicated by UpSampling*(1 / 2) in). After convolution processing, the filled depth map is output.

[0106] Step 5: Determine whether the water level exceeds the warning line:Figure 7 It is a flowchart for determining whether the water level exceeds the warning line according to an embodiment of the present invention. As Figure 7 shown, the starting point of the warning line and the mask image of the water area are input. First, 20 points are sampled from the warning line in equal proportion. Then, the pixel values corresponding to these 20 points in the mask image are counted. If the point is within the water area, the pixel value of this point is 255; if it is not within the water area, the pixel value is 0. Finally, if the number of points with pixel value equal to 255 among the 20 points exceeds 18 (accounting for 0.9), it can be considered that the water level has exceeded the warning line at this time, and an alarm is immediately sent to the platform.

[0107] Step 6, calculation of the relative distance between the water level line and the warning line: For the case where the water level line does not exceed the warning line, the relative distance between the water level line and the warning line needs to be calculated. In the past, the water level line was generally considered relatively simple and regarded as a straight line. The method of linear fitting was used to fit the water level line, and then the minimum distance between the two line segments (the warning line and the water level line fitted as a straight line) was calculated. However, in actual business scenarios, the water level line often has irregular shapes, including polygons and irregular curves. Calculating the minimum distance from a line segment on the image coordinate system to a closed irregular curve, calculating the distances between M points on the straight line and N points on the irregular curve, and finally taking the minimum distance. Taking the irregular curve water level line in a 1920*1080 image as an example, M is at least 50 and N is at least 4000. Calculating the distances one by one will bring a large computational overhead. The embodiment of the present invention proposes an algorithm for quickly estimating the minimum relative distance between the warning line and the water level line. Figure 8 It is a flowchart for calculating the relative distance between the water level line and the warning line according to an embodiment of the present invention. As Figure 8 shown, in order to reduce the computational amount, first downsample the mask image to 1 / 4 of the original image, and synchronously reduce the starting point coordinates of the warning line to 1 / 4 of the original. Then, judge the direction of the warning line in the water area, which may be in the up, down, left, or right direction. Then, 20 points are sampled from the downsampled warning line. If the warning line is above or below the water area, the intersection points of the sampled points of the warning line and the water area in the Y direction will be found. If it is on the left or right, the intersection points of the sampled points and the water area in the X direction will be found. After obtaining 20 intersection points, 5 points with the smallest distance from the corresponding sampled points are selected from the 20 intersection points as the initial points. The distances between each initial point and the 20 sampled points on the warning line are calculated respectively. To speed up the calculation, the binary search method is used here. Starting from the left and right endpoints of the sampled points, the minimum distance between each initial point and all the sampled points on the warning line is gradually calculated through binary search. Finally, the minimum distances of the 5 initial points from the sampled points on the warning line are compared, and the minimum value is selected. At this time, if the warning line is above or below the water area, when calculating the relative distance, it needs to be divided by the height of the downsampled mask image. If it is in the left or right direction, it is divided by the width of the downsampled mask image.

[0108] Through the above embodiments, the present invention can avoid the problem of misidentifying the water level line caused by incorrect water area segmentation, missed segmentation, or inaccurate segmentation, overcome the problem of poor generalization of conventional deep learning segmentation models, and can handle various scenarios.

[0109] Meanwhile, the embodiments of the present invention can also avoid the problem of misidentifying the water level line under occlusion, and can alarm when the water level line exceeds the warning line, and also give a quantitative description of the situation where the water level line does not exceed the warning line, which helps to give an early warning.

[0110] The following will be described in detail in conjunction with another embodiment.

[0111] Embodiment 2

[0112] A water level recognition device based on image fusion provided in this embodiment includes multiple implementation units, and each implementation unit corresponds to each implementation step in the first embodiment above. Its specific implementation manner and beneficial effects can be referred to the foregoing method embodiment, and will not be elaborated here.

[0113] Figure 9 is a schematic diagram of an optional water level recognition device based on image fusion according to an embodiment of the present invention. As Figure 9 shown, the water level recognition device based on image fusion may include: an image segmentation unit 91, a depth estimation unit 92, an image fusion unit 93, a water level line extraction unit 94, and a report output unit 95.

[0114] Among them, the image segmentation unit 91 is used to perform segmentation processing on the target image detecting the existence of water area by using a target instance segmentation model, and output a global contour segmentation map including the water area and an initial mask map for initially segmenting the water area.

[0115] The depth estimation unit 92 is used to perform depth estimation on the target image by using an improved target depth estimation model to generate a depth map.

[0116] The image fusion unit 93 is used to input the initial mask map, the global contour segmentation map, and the depth map into a multi-image fusion network model, and output a water area segmentation mask map for complementing the occluded water area part.

[0117] The water level line extraction unit 94 is used to extract the water level line for calibrating the water area mask region in the water area segmentation mask map.

[0118] The report output unit 95 is used to output a water level detection report based on the comparison result between the water level line and the warning line.

[0119] In this embodiment, the depth of the target image can be estimated by an improved target depth estimation model, the occluded water area can be complemented by a multi-image fusion network model, and then the water level line calibrating the water area mask region in the water area segmentation mask map is extracted and compared with the warning line to determine the water level position, solving the problems of mis-segmentation and missed segmentation existing in the existing water area segmentation algorithms, improving the accuracy of water level position recognition, and thus solving the technical problem of low water level detection accuracy due to object occlusion and mis-segmentation of the water area in the related art.

[0120] Optionally, the target instance segmentation model is obtained in the following manner: an initial vision model based on the Transformer architecture is constructed by multiple recognition modules introducing a spatial attention mechanism; the initial vision model is used to distill a preset image classification model, and the segmentation parameters of the preset image classification model for the entity objects in the image are adjusted to obtain the target instance segmentation model.

[0121] Optionally, the image segmentation unit includes: a global segmentation module for globally segmenting the target image by using the target instance segmentation model to obtain a global contour segmentation map including the water area; a water area segmentation module for using the water area target detection frame as a prompt feature and segmenting the water area of the target image by using the target instance segmentation model to obtain an initial mask map for initially segmenting the water area.

[0122] Optionally, the target depth estimation model is obtained in the following manner: the target image encoder is adjusted by using the unlabeled water area dataset of the self-service scenario; the target image encoder is frozen, and the initial depth estimation model is adjusted by using the labeled open-source depth estimation dataset; the adjusted initial depth estimation model is used to infer the unlabeled water area data of the self-service scenario to obtain a depth estimation map; a preset data transformation strategy is used to transform the obtained depth estimation map to obtain a labeled water area dataset; the labeled water area dataset is used to train the specified image encoder in the adjusted preset image classification model to obtain the target depth estimation model.

[0123] Optionally, the water level recognition device based on image fusion further includes: a water area detection unit for detecting whether there is a water area in the target image by using a preset target detection model before segmenting the target image with a detected water area by using the target instance segmentation model, where the target detection model is obtained by modifying multiple feature hierarchical structures of the object detection algorithm based on deep learning, and the modification of the multiple feature hierarchical structures includes: deleting the targets in the image smaller than the preset target size and retaining the targets in the image larger than the preset target size.

[0124] Optionally, the water level recognition device based on image fusion further includes: a channel splicing unit, configured to perform channel splicing processing on the initial mask map, the global contour segmentation map, and the depth map by a multi-image fusion network model after inputting the initial mask map, the global contour segmentation map, and the depth map into the multi-image fusion network model; a convolution unit, configured to perform multiple convolution processes on the spliced image by the multi-image fusion network model to output a water area segmentation mask map for complementing the occluded water area, wherein, during the convolution process, a channel attention mechanism and a spatial attention mechanism are introduced. The channel attention mechanism is used to extract the global information of each channel of the image through global average pooling processing and global maximum pooling processing, and then generate the weights of the channels through multiple activation functions. The weights are used to weight each channel of the input feature map to enhance the features of the target channel; the spatial attention mechanism is used to perform global maximum pooling processing and global average pooling processing on the spliced image in the channel dimension to generate a spatial attention map, and the spatial attention map is used to identify the spatial positions in the spliced image and enhance the spatial positions. Wherein, during the convolution process of the multi-image fusion network model on the image, upsampling processing and downsampling processing are performed on the image.

[0125] Optionally, the multi-image fusion network model is obtained in the following manner: constructing a synthetic occlusion dataset and a real occlusion dataset, wherein the synthetic occlusion dataset refers to pasting a mask of a water area carrying a specified feature object on the edge of an unoccluded water area, providing labels for the unoccluded data, and the specified feature object includes at least: trees, rocks, and moss; training a pre-constructed convolutional neural network architecture based on image segmentation using the synthetic occlusion dataset and the real occlusion dataset to obtain the multi-image fusion network model.

[0126] Optionally, the report output unit includes: a mask map input unit, configured to input the starting point of the warning line and the water area segmentation mask map; a feature point sampling unit, configured to sample a plurality of feature points from the warning line in equal proportion and count the pixel values corresponding to the plurality of feature points in the water area segmentation mask map; a first confirmation unit, configured to confirm that the water level exceeds the warning line and record the water level line and the warning information into the water level detection report when the number of pixel values indicating that the feature points are located in the water area exceeds a preset quantity threshold; a second confirmation unit, configured to confirm that the water level does not exceed the warning line and record the water level line into the water level detection report when the number of pixel values indicating that the feature points are located in the water area does not exceed the preset quantity threshold.

[0127] Optionally, the water level recognition device based on image fusion further includes: a downsampling unit, configured to perform downsampling processing on the water area segmentation mask map and reduce the starting point coordinates of the warning line after outputting a water level detection report based on the comparison result between the water level line and the warning line; a direction detection unit, configured to detect the direction of the warning line in the water area and perform feature sampling on the reduced warning line; a first selection unit, configured to select the intersection points of the sampling points of the warning line and the water area in the Y direction when the warning line is above and below the water area; a second selection unit, configured to select the intersection points of the sampling points of the warning line and the water area in the X direction when the warning line is on the left and right sides of the water area; an initial point selection unit, configured to select a predetermined number of intersection points from all the selected intersection points and use the selected intersection points as the initial points, where the selected initial points are the intersection points with the smallest distance from the corresponding sampling points; a distance value determination unit, configured to calculate the distance values between each initial point and each sampling point on the warning line, and determine the minimum distance value between all the sampling points on the warning line and the initial point through a binary search strategy; a warning distance value selection unit, configured to compare the minimum distance values between all the initial points and the sampling points on the warning line and select the minimum distance value as the target warning distance value; a first relative distance value determination unit, configured to divide the target warning distance value by the height value of the water area segmentation mask map to obtain a relative distance if the warning line is above and below the water area; or a second relative distance value determination unit, configured to divide the target warning distance value by the width value of the water area segmentation mask map to obtain a relative distance if the warning line is on the left and right sides of the water area, where the relative distance is used to record the water level position change information.

[0128] The above-mentioned water level recognition device based on image fusion may further include a processor and a memory. The above-mentioned image segmentation unit 91, depth estimation unit 92, image fusion unit 93, water level line extraction unit 94, report output unit 95, etc. are all stored in the memory as program units, and the processor executes the above program units stored in the memory to implement corresponding functions.

[0129] The above-mentioned processor includes a kernel, and the kernel retrieves the corresponding program units from the memory. One or more kernels can be set, and by adjusting the kernel parameters, the water level recognition based on the multi-image fusion technology can send the comparison result to the target terminal.

[0130] The above-mentioned memory may include non-permanent memory in a computer-readable medium, random access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash memory (flash RAM), and the memory includes at least one storage chip.

[0131] Embodiment 3

[0132] Embodiments of the present application can provide an electronic deviceFigure 10 is a structural block diagram of an electronic device according to an embodiment of the present application. As Figure 10 shown, the electronic device may include: one or more ( Figure 10 only one is shown in the figure) processors 1002, a memory 1004, a storage controller, and a peripheral interface, where the peripheral interface is connected to a radio frequency module, an audio module, and a display.

[0133] Among them, the memory can be used to store software programs and modules, such as the program instructions / modules corresponding to the water level recognition method and device based on image fusion in the embodiments of the present application. The processor executes various functional applications and data processing by running the software programs and modules stored in the memory, that is, implements the above-mentioned water level recognition method based on image fusion. The memory may include a high-speed random access memory, and may also include a non-volatile memory, such as one or more magnetic storage devices, flash memories, or other non-volatile solid-state memories. In some instances, the memory may further include a memory remotely provided relative to the processor, and these remote memories can be connected to the terminal through a network. Examples of the above network include but are not limited to the Internet, enterprise intranets, local area networks, mobile communication networks, and combinations thereof.

[0134] The processor can call the information and application programs stored in the memory through a transmission device to perform the following steps: segmenting a target image detected with a water area using a target instance segmentation model, and outputting a global contour segmentation map including the water area and an initial mask map for a preliminary segmentation of the water area; estimating the depth of the target image through an improved target depth estimation model to generate a depth map; inputting the initial mask map, the global contour segmentation map, and the depth map into a multi-image fusion network model, and outputting a water area segmentation mask map for complementing the occluded water area part; extracting a water level line for calibrating the water area mask region in the water area segmentation mask map; and outputting a water level detection report based on the comparison result between the water level line and the warning line.

[0135] Those of ordinary skill in the art can understand that Figure 10 the structure shown is only schematic, and the electronic device can also be a terminal device such as a smart phone, a tablet computer, a handheld computer, and a Mobile Internet Device (MID), a PAD, etc. Figure 10 It does not limit the structure of the above-mentioned electronic device. For example, the electronic device may further include more or fewer components (such as a network interface, a display device, etc.) than those shown in Figure 10 the figure, or have a different configuration from that shown in Figure 10 the figure.

[0136] Those of ordinary skill in the art can understand that all or part of the steps in the above-described various image fusion-based water level recognition methods can be completed by instructing the relevant hardware of the terminal device through a program, and this program can be stored in a computer-readable storage medium. The storage medium can include: a flash drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disc, etc.

[0137] Embodiment 4

[0138] An embodiment of the present application further provides a storage medium. Optionally, in this embodiment, the above storage medium can be used to store the program code executed by the image fusion-based water level recognition method provided in the first embodiment above.

[0139] According to another aspect of the embodiments of the present invention, there is also provided a computer-readable storage medium. The computer-readable storage medium includes a stored computer program, wherein when the computer program runs, it controls the device where the computer-readable storage medium is located to execute any one of the image fusion-based water level recognition methods in the first embodiment above.

[0140] Optionally, in this embodiment, the above storage medium can be located in any one of the computer terminals in a computer terminal group in a computer network, or in any one of the mobile terminals in a mobile terminal group.

[0141] The present application also provides a computer program product, including a computer program, and the steps of the image fusion-based water level recognition method described in various embodiments of the present application are implemented when the computer program is executed by a processor.

[0142] The present application also provides a computer program product, including a non-volatile computer-readable storage medium. The non-volatile computer-readable storage medium stores a computer program, and the steps of the image fusion-based water level recognition method described in various embodiments of the present application are implemented when the computer program is executed by a processor.

[0143] The above serial numbers of the embodiments of the present invention are only for description and do not represent the advantages or disadvantages of the embodiments.

[0144] In the above embodiments of the present invention, the descriptions of each embodiment have their own emphases. For the parts not detailed in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.

[0145] In several embodiments provided by the present application, it should be understood that the disclosed technical content can be implemented in other ways. Among them, the device embodiments described above are merely illustrative. For example, the division of the units can be a logical function division. In actual implementation, there can be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the couplings or direct couplings or communication connections shown or discussed with each other can be through some interfaces. The indirect couplings or communication connections of units or modules can be in electrical or other forms.

[0146] The units described as separate components may or may not be physically separated. The components shown as units may or may not be physical units, that is, they can be located in one place or distributed to multiple units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0147] In addition, in each embodiment of the present invention, the functional units can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above-mentioned integrated units can be implemented in the form of hardware or in the form of software functional units.

[0148] If the above-mentioned integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server or a network device, etc.) to execute all or part of the steps of the methods described in each embodiment of the present invention. And the aforementioned storage medium includes: USB flash drives, read-only memories (ROMs), random access memories (RAMs), mobile hard disks, magnetic disks or optical discs and other various media that can store program codes.

[0149] The above are only the preferred embodiments of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present invention, several improvements and refinements can still be made, and these improvements and refinements should also be regarded as the protection scope of the present invention.

Claims

1. A water level recognition method based on image fusion, characterized in that, Including: Using a target instance segmentation model to perform segmentation processing on a target image detected with water areas, and outputting a global contour segmentation map including the water areas and an initial mask map that preliminarily segments the water areas; Performing depth estimation on the target image through an improved target depth estimation model to generate a depth map; Inputting the initial mask map, the global contour segmentation map, and the depth map into a multi-image fusion network model, and outputting a water area segmentation mask map that complements the occluded water area part; Extracting a water level line that calibrates the water mask area in the water area segmentation mask map; Based on the comparison result between the water level line and the warning line, outputting a water level detection report.

2. The water level identification method according to claim 1, characterized in that The target instance segmentation model is obtained through the following method: Constructing an initial vision model based on the Transformer architecture through multiple recognition modules introducing a spatial attention mechanism; Using the initial vision model to distill a preset image classification model, and adjusting the segmentation parameters of the preset image classification model for entity objects in the image to obtain the target instance segmentation model.

3. The water level identification method according to claim 2, characterized in that, The step of using a target instance segmentation model to perform segmentation processing on a target image detected with water areas and outputting a global contour segmentation map including the water areas and an initial mask map that preliminarily segments the water areas includes: Performing global segmentation on the target image using the target instance segmentation model to obtain the global contour segmentation map including the water areas; Using the water area target detection box as a prompt feature, and performing water area segmentation on the target image using the target instance segmentation model to obtain the initial mask map that preliminarily segments the water areas.

4. The water level identification method according to claim 2, characterized in that The target depth estimation model is obtained through the following method: Adjusting a target image encoder using an unlabeled water area dataset from the self-service scenario; Freezing the target image encoder, and using a labeled open-source depth estimation dataset to adjust an initial depth estimation model; Using the adjusted initial depth estimation model to infer the unlabeled water area data from the self-service scenario to obtain a depth estimation map; Performing transformation on the obtained depth estimation map using a preset data transformation strategy to obtain a labeled water area dataset; Training a specified image encoder in the adjusted preset image classification model using the labeled water area dataset to obtain the target depth estimation model.

5. The water level recognition method according to claim 1, characterized in that, Before using a target instance segmentation model to perform segmentation processing on a target image detected with water areas, it further includes: Detecting whether there is a water area in the target image using a preset target detection model, where the target detection model is obtained by modifying multiple feature hierarchical structures of a deep learning-based target detection algorithm. Modifying the multiple feature hierarchical structures includes: deleting targets in the image smaller than a preset target size and retaining targets in the image larger than the preset target size.

6. The water level recognition method according to claim 1, characterized in that, After inputting the initial mask map, the global contour segmentation map, and the depth map into a multi-image fusion network model, it further includes: The multi-image fusion network model performs channel splicing processing on the initial mask map, the global contour segmentation map, and the depth map; The spliced image is subjected to multiple convolutional processes by the multi-image fusion network model to output a water area segmentation mask map for complementing the occluded water area. Among them, during the convolutional process, a channel attention mechanism and a spatial attention mechanism are introduced. The channel attention mechanism is used to extract the global information of each channel of the image through global average pooling and global maximum pooling processes, and then generate the weights of the channels through multiple activation functions. The weights are used to weight each channel of the input feature map to enhance the features of the target channel. The spatial attention mechanism is used to perform global maximum pooling and global average pooling processes on the spliced image in the channel dimension to generate a spatial attention map, and the spatial attention map is used to identify the spatial positions in the spliced image and enhance the spatial positions. Among them, during the convolutional process of the multi-image fusion network model on the image, upsampling and downsampling processes are performed on the image.

7. The water level recognition method according to claim 6, characterized in that The multi-image fusion network model is obtained through the following method: Construct a synthetic occlusion dataset and a real occlusion dataset. Among them, the synthetic occlusion dataset refers to pasting a mask of a water area carrying a specified feature object on the edge of an unoccluded water area and providing labels for the unoccluded data. The specified feature objects at least include: trees, rocks, and moss. Use the synthetic occlusion dataset and the real occlusion dataset to train a pre-constructed convolutional neural network architecture based on image segmentation to obtain the multi-image fusion network model.

8. The water level recognition method according to claim 1, characterized in that The steps of outputting a water level detection report based on the comparison result between the water level line and the warning line include: Input the starting point of the warning line and the water area segmentation mask map. Sample multiple feature points from the warning line in equal proportion and count the pixel values corresponding to the multiple feature points in the water area segmentation mask map. In the case where the number of pixel values indicating that the feature points are located within the water area exceeds a preset number threshold, confirm that the water level exceeds the warning line, and record the water level line and the warning information into the water level detection report. In the case where the number of pixel values indicating that the feature points are located within the water area does not exceed the preset number threshold, confirm that the water level does not exceed the warning line, and record the water level line into the water level detection report.

9. The water level identification method according to claim 1, wherein After outputting the water level detection report based on the comparison result between the water level line and the warning line, it further includes: Perform downsampling on the water area segmentation mask map and reduce the starting point coordinates of the warning line. Detect the direction of the warning line in the water area and perform feature sampling on the reduced warning line. If the warning line is above and below the water area, select the intersection points of the sampling points of the warning line and the water area in the Y direction. If the warning line is on the left and right of the water area, select the intersection points of the sampling points of the warning line and the water area in the X direction. Select a predetermined number of intersection points from all the selected intersection points and use the selected intersection points as the initial points, where the selected initial points are the intersection points with the smallest distance from the corresponding sampling points. For each of the initial points, calculate the distance values between the initial point and each sampling point on the warning line, and determine the minimum distance value between the initial point and all the sampling points on the warning line through a binary search strategy; Compare the minimum distance values between the initial points and the sampling points on the warning line, and select the minimum distance value as the target warning distance value; If the warning line is above and below the water area, divide the target warning distance value by the height value of the water area segmentation mask map to obtain a relative distance; or, If the warning line is on the left and right of the water area, divide the target warning distance value by the width value of the water area segmentation mask map to obtain a relative distance, where the relative distance is used to record the water level position change information.

10. A water level recognition device based on image fusion, characterized in that, Comprising: An image segmentation unit for segmenting a target image with detected water areas by using a target instance segmentation model, and outputting a global contour segmentation map including the water areas and an initial mask map for preliminary segmentation of the water areas; A depth estimation unit for estimating the depth of the target image by using an improved target depth estimation model to generate a depth map; An image fusion unit for inputting the initial mask map, the global contour segmentation map, and the depth map into a multi-image fusion network model to output a water area segmentation mask map for complementing the occluded water area part; A water level line extraction unit for extracting a water level line for calibrating the water area mask region in the water area segmentation mask map; A report output unit for outputting a water level detection report based on the comparison result between the water level line and the warning line.

11. An electronic device, characterized in that, Comprising one or more processors and a memory, where the memory is used to store one or more programs, and when the one or more programs are executed by the one or more processors, the one or more processors implement the image fusion-based water level recognition method according to any one of claims 1 to 9.

12. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, the steps of the image fusion-based water level recognition method according to any one of claims 1 to 9 are implemented.

Citation Information

Cited By

  • Body posture recognition method and system, intelligent terminal and storage medium

    CN121281099A

  • A human posture recognition method, system, intelligent terminal and storage medium

    CN121281099B

  • Image feature extraction method and device, equipment and storage medium

    CN121415083A