Flood or stagnant water detection device and method

The flood or standing water detection device employs neural networks with attention mechanisms for semantic segmentation to enhance detection accuracy by learning contextual information, addressing the limitations of manual monitoring and improving classification precision.

JP7893074B2Active Publication Date: 2026-07-22FUJITSU LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
FUJITSU LTD
Filing Date
2022-07-13
Publication Date
2026-07-22

AI Technical Summary

Technical Problem

Current methods for monitoring river floods and road ponding rely on manual observation, which is inadequate in areas where visibility is limited, leading to traffic disruptions and vehicle submersion due to delayed detection of floods and ponding.

Method used

A flood or standing water detection device that utilizes a neural network model with attention mechanisms for semantic segmentation, dividing image data into semantic classes based on pixel correlations and contextual information to calculate water occupancy rates and determine flood or standing water presence.

Benefits of technology

Improves detection accuracy by learning important contextual information and mitigating the impact of imbalanced datasets, enhancing classification accuracy and flood or stagnant water detection precision.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007893074000001
    Figure 0007893074000001
  • Figure 0007893074000002
    Figure 0007893074000002
  • Figure 0007893074000003
    Figure 0007893074000003
Patent Text Reader

Abstract

To provide a detector and a method for detecting a flood or stored water according to an embodiment.SOLUTION: The device includes: a division unit for dividing each frame image data of a video stream to different meaning classes on the basis of the correlation between each pixel and all the positional pixels in each frame image data of the video steam and / or context information in the image data of each frame of the video stream; a calculation unit for calculating the occupation rate of water in at least one interest region on the basis of the result of division; and a determination unit for detecting the presence or absence of a flood or stored water on the basis of the occupation rate of water in at least one interest region.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of information technology.

Background Art

[0002] In recent years, many disasters have occurred due to abnormal weather, and floods are one of the most serious disasters. The occurrence of flood disasters not only causes deaths of people due to accidents such as drowning and electric shock, but also causes ponding on roads, greatly hindering the development of urban traffic and urbanization. Therefore, the monitoring and prediction of river floods and ponding on roads are of great value and significance.

[0003] Note that the description of the above background art is for the purpose of clearly and completely understanding the technical solution of the present invention and is described to enable those skilled in the art to understand. These technical solutions are only described as the background art part of the present invention and are not well-known to those skilled in the art.

Summary of the Invention

Problems to be Solved by the Invention

[0004] Currently, generally, the monitoring methods for river floods and ponding on roads mainly rely on manual monitoring on-site. However, in some areas, it is difficult to observe the river or ponding on roads, so floods and ponding cannot be observed immediately. Therefore, since drivers cannot obtain information about ponding, driving into the area of ponding may cause traffic jams and vehicles may be submerged.

[0005] To solve at least one of the above problems, embodiments of the present invention provide a detection device and method for floods or ponding.

Means for Solving the Problems

[0006] In a first embodiment of the present invention, a flood or standing water detection device is provided, comprising: a division unit that divides each frame image data of a video stream into different semantic classes based on the correlation between each pixel in each frame image data of a video stream and all position pixels, and / or context information in the image data of each frame of the video stream; a calculation unit that calculates the water occupancy rate in at least one region of interest based on the division results; and a determination unit that detects whether or not a flood or standing water has occurred based on the water occupancy rate in at least one region of interest.

[0007] A second embodiment of the present invention provides a method for detecting a flood or standing water, comprising the steps of: dividing each frame image data of a video stream into different semantic classes based on the correlation between each pixel in each frame image data of a video stream and all position pixels, and / or contextual information in each frame image data of the video stream; calculating the percentage of water occupancy in at least one region of interest based on the division results; and detecting whether a flood or standing water has occurred based on the percentage of water occupancy in at least one region of interest.

[0008] The advantageous effects of the embodiments of the present invention are as follows: By considering contextual information and / or correlations between each position pixel during semantic segmentation, more important information can be learned, better feature vectors for semantic segmentation can be extracted, the negative impact of unbalanced training datasets on semantic segmentation can be mitigated, classification accuracy can be improved, and the accuracy of flood or stagnant water detection can be improved.

[0009] Specific embodiments of the present invention are disclosed in detail as shown in the following description and drawings, illustrating ways in which the principles of the present invention can be employed. However, the embodiments of the present invention are not limited in scope. Embodiments of the present invention include various modifications, alterations, and equivalents within the spirit and content of the appended claims.

[0010] Features described and / or shown in one embodiment may be used in the same or similar manner in one or more other embodiments, may be combined with features in other embodiments, or may substitute for features in other embodiments.

[0011] When used in this text, the terms "contain / have" mean the presence of a feature, element, step, or constituent element, and do not exclude the presence or addition of one or more other features, elements, steps, or constituent elements. [Brief explanation of the drawing]

[0012] The drawings included herein are for illustrative purposes to illustrate embodiments of the present invention and constitute part of this specification, illustrating embodiments of the present invention together with the written description to explain the principles of the present invention. It should be noted that the drawings described herein are merely for illustrating embodiments of the present invention, and those skilled in the art can readily obtain other drawings based on these. [Figure 1] This is a schematic diagram of one flood or stagnant water detection device according to Embodiment 1 of the present invention. [Figure 2A] This is a schematic diagram of the attention module's configuration. [Figure 2B] This is a schematic diagram of the attention module's configuration. [Figure 3] This is a schematic diagram of one configuration of a neural network model according to Embodiment 1 of the present invention. [Figure 4] This is a schematic diagram of one configuration of a divided section according to Embodiment 1 of the present invention. [Figure 5A] This is a schematic diagram of the original image according to Example 1 of the present invention. [Figure 5B] It is a schematic diagram of a pseudo-image corresponding to the original image according to Example 1 of the present invention. [Figure 6A] It is a schematic diagram of the original image according to Example 1 of the present invention. [Figure 6B] It is a schematic diagram of a pseudo-image corresponding to the original image according to Example 1 of the present invention. [Figure 7A] It is a schematic diagram of the original image according to Example 1 of the present invention. [Figure 7B] It is a schematic diagram of a pseudo-image corresponding to the original image according to Example 1 of the present invention. [Figure 10A] It is a schematic diagram of the region of interest of the original image when no flood occurs according to Example 1 of the present invention. [Figure 10B] It is a schematic diagram of the region of interest of the original image when a flood occurs according to Example 1 of the present invention. [Figure 11] It is a schematic diagram of the region of interest of the original image when a flood occurs on a road according to Example 1 of the present invention. [Figure 12] It is a schematic diagram of the region of interest of the original image when a flood occurs on a road according to Example 1 of the present invention. [Figure 13] It is a schematic diagram of one electronic device according to Example 2 of the present invention. [Figure 14] It is a schematic block diagram of the system configuration of the electronic device according to Example 2 of the present invention. [Figure 15] It is a schematic diagram of one method for detecting flood or standing water according to Example 3 of the present invention. It is a schematic diagram of one aspect of step 1301 of the embodiment of FIG. 13. It is a schematic diagram of the method for detecting flood or standing water according to Example 3 of the present invention.

[0013]

Mode for Carrying Out the Invention

[0014] The above and other features of the present invention can be understood from the drawings and the following description. In the specification and the drawings, specific embodiments of the present invention, that is, some embodiments in accordance with the principles of the present invention are disclosed. It should be noted that the present invention is not limited to the described embodiments, and the present invention includes all modifications, variations, and equivalents within the scope of the claims.

[0014] <Example 1> An embodiment of the present invention provides a flood or ponding water detection device. FIG. 1 is a schematic diagram of one of the flood or ponding water detection devices according to Example 1 of the present invention.

[0015] As shown in FIG. 1, the flood or ponding water detection device 100 includes a segmentation unit 101, a calculation unit 102, and a determination unit 103.

[0016] The segmentation unit 101 divides each frame image data of the video stream into different semantic classes based on the correlation relationship between each pixel and all position pixels in each frame image data of the video stream, and / or the context information in the image data of each frame of the video stream.

[0017] <I The calculation unit 102 calculates the occupancy rate of water in at least one region of interest based on the segmentation result.

[0018] The determination unit 103 detects whether flood or ponding water has occurred based on the occupancy rate of water in at least one region of interest.

[0019] In an embodiment of the present invention, each frame image data of the video stream processed by the flood or ponding water detection device 100 may be each frame image data of a real-time video stream obtained from an imaging device that captures an area where a lane is located or an area where a river is located. For example, it may be each frame image data of a real-time video stream captured by a surveillance camera installed above a road or a surveillance camera installed near a river.

[0020] In some embodiments, each frame image data may be a pre-processed image, which may include adjusting brightness and / or cropping size, but embodiments of the present invention are not limited thereto.

[0021] In some embodiments, the segmentation unit 101 may perform semantic segmentation on image data using a trained neural network model. Here, the training set for training the neural network includes a river training set and a road training set. The river training set includes a training set for cases where no flood occurs and a training set for cases where flood occurs. The main semantic classes (categories) of each image containing a river include river, sky, riverbank, riverbed, bridge, vegetation, water facilities and bridges, etc. Therefore, the datasets of embodiments of the present invention include only these semantic classes. The road training set includes a training set for cases where no standing water occurs and a training set for cases where standing water occurs on the road. The main semantic classes of each image containing a road further include automobile, bicycle, pedestrian, road, sidewalk, fence, utility pole, streetlamp, etc., other examples are omitted here.

[0022] In some embodiments, the neural network model may be ResNet50, InceptionV3, Xception model, etc., and details thereof may be found in related technologies, and embodiments of the present invention are not limited to these. The dilated Xception65 model is described below as an example.

[0023] In some embodiments, the splitting unit 101 may divide the image data of each frame of the video stream into different semantic classes based on the correlation between each pixel in the image data of each frame of the video stream and all the position pixels. In other words, the neural network model described above learns more important information using an attention mechanism (by adding an attention module). This attention includes positional attention and / or channel attention.

[0024] In some embodiments, all of the above position pixels include pixels at each position on the same image data and pixels at each position on different channels. The correlation between each pixel point in the image data of each frame of the video stream and all the position pixels may be determined by position attention and / or channel attention.

[0025] Figures 2A and 2B are schematic diagrams of the configuration of the attention module. Figure 2A is a schematic diagram of the module for calculating the feature vector of the position attention. As shown in Figure 2A, in the case of position attention, image data is input to a neural network model (using a dilated Xception65 as the backbone) to extract a feature vector matrix (represented as a C×H×W dimensional matrix), and a position attention matrix (represented as an (H×W)×(H×W) dimensional matrix) is generated based on this feature vector matrix. This position attention matrix can represent the correlation between any two pixels in the feature vector. By performing a convolution multiplication on the position attention matrix and the original feature vector matrix, and then adding them together, a feature vector matrix that takes positional correlations into account (represented as a C×H×W dimensional matrix) is obtained. The original feature vector matrix may be reshaped into a matrix C×(H×W). The position attention matrix is ​​obtained by transposing the matrix C×(H×W) and then multiplying it by the matrix C×(H×W). Preferably, the position attention matrix may be processed using the softmax function, convolved with the reshaped original feature vector matrix, and then added to the original feature vector matrix.

[0026] Figure 2B is a schematic diagram of the configuration of the feature vector calculation module for channel attention. As shown in Figure 2B, in the case of channel attention, image data is input to a neural network model (using a dilated Xception65 as the backbone) to extract a feature vector matrix (represented as a C×H×W dimensional matrix), and a channel attention matrix (represented as a C×C dimensional matrix) is generated based on this feature vector matrix. This channel attention matrix can represent the correlation between any two pixels between channels. A feature vector matrix that considers the channel correlation (represented as a C×H×W dimensional matrix) is obtained by performing a convolution multiplication on the original feature vector matrix and the channel attention matrix, and then adding them together. The original feature vector matrix may be reshaped into a matrix C×(H×W). The channel attention matrix is ​​obtained by transposing the matrix C×(H×W) and then multiplying it by the matrix C×(H×W). Preferably, the channel attention matrix may be processed using the softmax function, convolved with the original feature vector matrix after reshaping, and then added back together.

[0027] In some embodiments, the feature vector matrices processed by channel attention and position attention may be merged (for example, by adding or cascading them and then using a convolution operation to recover the number of channels) to obtain feature vectors for semantic segmentation.

[0028] According to the above embodiment, by adding an attention mechanism to the neural network model, it is possible to learn correlations between distant pixels (long-range semantic information), thereby enabling the extraction of better feature vectors for semantic segmentation.

[0029] In some embodiments, the segmentation unit 101 may divide the image data of each frame of the video stream into different semantic classes based on contextual information in the image data of each frame of the video stream. In other words, semantic segmentation can be more effectively achieved by learning more important information using contextual information (by adding a context module) to the neural network model described above.

[0030] For example, context information may be obtained using a context-priority model. This context information includes correlations between pixels within the same class and correlations between pixels within different classes. This context-priority model may be the Context Prior Layer in a CPNet model. This model uses an affinity loss function, the details of which can be found in related technologies and are omitted here.

[0031] In some embodiments, the two embodiments of the segmentation unit 101 may be implemented independently or in combination. Figure 3 is a schematic diagram of the configuration of the neural network model when implemented in combination. As shown in Figure 3, when implemented in combination, feature vectors are extracted using Xception 65 as the backbone, the extracted feature vectors are processed by an attention mechanism, the processed feature vectors are used as input to a context-aware model, the output of the context-aware model is classified (input to a prediction network), and semantic segmentation results are obtained.

[0032] Figure 4 is a schematic diagram of the configuration of one embodiment of the divided section 101. As shown in Figure 4, the divided section 101 includes the following modules.

[0033] The extraction module 201 inputs each frame image data into the feature extraction network and extracts a feature map.

[0034] The decision module 202 determines the correlation between any two locations on the feature map.

[0035] The update module 203 determines weights based on the correlation and updates the features at each location in the feature map based on these weights.

[0036] The segmentation module 204 extracts context information from the updated features, inputs the context information and the updated features into the prediction network, and obtains the segmentation result of the semantic class for each frame image data.

[0037] In some embodiments, the extraction module 201 uses, for example, the dilated Xception65 described above as a feature extraction network to extract a feature map. The decision module 202 and the update module 203 update the features at each position on the feature map using the attention module described above. The position attention matrix and / or channel attention matrix described above may represent the correlation, and the convolution result of the attention matrix and the reshaped original feature map may be considered as the weight. The specific processing processes of the decision module 202 and the update module 203 can be seen in Figures 2A and 2B, and are omitted here. The partition module 204 extracts contextual information using, for example, the context prior model described above, such as a Context Prior Layer, and inputs the output of the context prior model into the prediction network to obtain the semantic classification result for each frame image data.

[0038] In some embodiments, after obtaining the semantic class partitioning results, the results of different classes may be represented by different colors, and after color mapping, pseudo-images may be generated. Figures 5A, 6A, and 7A are schematic diagrams of the original images, and Figures 5B, 6B, and 7B are schematic diagrams of the corresponding pseudo-images.

[0039] In some embodiments, after obtaining the semantic class partitioning results, the following processing may be performed on the partitioning results, and then the water occupancy rate in at least one region of interest may be calculated.

[0040] In some embodiments, the apparatus further includes a second processing unit (optional, not shown).

[0041] The second processing unit performs noise reduction processing on the division results.

[0042] For example, the second processing unit may perform noise reduction processing by methods such as eroding and dilating, but the embodiments of the present invention are not limited thereto. The method of eroding and dilating may refer to the prior art, and its explanation is omitted here.

[0043] In some embodiments, when detecting water in a road area, the image data may contain many moving objects, which may obstruct the road or puddle area and affect the detection results. For this reason, the device further includes a background extraction unit (optional, not shown).

[0044] The background extraction unit determines moving objects in the image based on a second number of frame image data and removes these moving objects from the segmentation result.

[0045] For example, the background extraction unit may determine moving objects in the image using methods such as interframe difference, background difference, or the ViBe algorithm. Details may be found in related technologies, and the embodiments of the present invention are not limited thereto. After determining the moving objects, the moving objects are removed from the segmentation result, or the pixel class corresponding to the moving object is modified to the semantic class to which the background belongs. For example, the pixel class at the position corresponding to the moving object (e.g., a vehicle) in the segmentation result is modified to a road to avoid the moving object obscuring a road or puddle area and to improve the accuracy of detection.

[0046] In some embodiments, the second number of frame images may be a second number of interval frames in the video stream. Interval frames mean multiple frames separated from each other by at least one frame, i.e., discontinuous frames. In this way, moving objects and stationary objects can be clearly distinguished, and moving objects can be easily identified.

[0047] In some embodiments, the apparatus further includes a first processing unit (optional, not shown).

[0048] The first processing unit counts the semantic class with the highest frequency of occurrence among the semantic classes of the same position in the first number of frame image data divided by the division unit 101, and sets the semantic class with the highest frequency of occurrence as the semantic class of that position.

[0049] For example, the division unit 101 determines a semantic class for each pixel in each of a first number of frames (e.g., consecutive frames). Here, the first processing unit sets the class of the pixel to the class that appears most frequently (most often) in the semantic segmentation results of the pixel in the first number of frames.

[0050] For example, for 10 (the first number) frame image data, the division unit 101 performs semantic segmentation on each, then scans the semantic segmentation results pixel by pixel to determine the class of each pixel in each frame, for example, determining whether the pixel belongs to a road or a vehicle. For a given pixel, if, as a result of counting across 10 frames, the first processing unit determines that the semantic class of the pixel is a road if 8 frames are classified as belonging to a road and 2 frames are classified as belonging to a vehicle, then the first processing unit determines that the pixel's semantic class is a road.

[0051] In some embodiments, the apparatus further includes a third processing unit (optional, not shown).

[0052] The third processing unit, when detecting a flood, determines a water mask and a riverbed mask based on the division results, and when detecting stagnant water, determines a water mask and a road mask based on the division results.

[0053] For the sake of explanation, the water mask will be referred to as mask1, the riverbed mask as mask2, and the road mask as mask3. Here, the size of each mask is the same as the size of each frame image data. In mask1, based on the division result, the value of pixels whose semantic class is water is set to 1, and the value of pixels at other locations is set to 0. Similarly, in mask2, based on the division result, the value of pixels whose semantic class is riverbed is set to 1, and the value of pixels at other locations is set to 0. In mask3, based on the division result, the value of pixels whose semantic class is road is set to 1, and the value of pixels at other locations is set to 0.

[0054] The following explains how to determine whether or not a flood has occurred.

[0055] In some embodiments, the calculation unit 102 calculates the ratio of the number of pixels of water in each of at least two regions of interest to the number of pixels in each region of interest, and the determination unit 103 determines whether or not a flood has occurred based on the comparison result of the ratio of water in each region of interest with a threshold corresponding to each region of interest.

[0056] In some embodiments, the at least two regions of interest include at least one primary region of interest (primary ROI) and at least one secondary region of interest (secondary ROI). Here, when detecting a flood, the primary region of interest includes the river region, or the river region and the riverbed region. The secondary region of interest includes the riverbank region and other areas where flooding may occur. Each of the above regions of interest is pre-configured before flood detection. Figures 8 and 9 are schematic diagrams of the pre-configured regions of interest in the original image. As shown in Figures 8 and 9, the image data includes one primary region of interest and three secondary regions of interest, the primary region of interest includes the river region, and the secondary regions of interest include the riverbank region.

[0057] In some embodiments, the calculation unit 102 may determine the ratio of the number of pixels in the water region to the number of pixels in the region of interest (which may be considered as the ratio of the area of ​​the water region to the area of ​​the region of interest) based on the water mask, or the water mask and the riverbed mask, and each region of interest. For example, let N1 be the number of pixels with a pixel value of 1 in the primary region of interest in mask1 or mask1 and mask2, and let Nz be the number of pixels in the primary region of interest. Let N2 be the number of pixels with a pixel value of 1 in one secondary region of interest in mask1 or mask1 and mask2, and let Nf be the number of pixels in that secondary region of interest. For the primary region of interest, the above ratio is N1 / Nz, and for one secondary region of interest, the above ratio is N2 / Nf. If there are other secondary regions of interest, the same method is used to determine the ratio for each secondary region of interest, and the explanation is omitted here. N1, Nz, N2, and Nf are integers of 1 or more.

[0058] In some embodiments, the determination unit 103 determines whether a flood has occurred based on a comparison between the ratio of water in each region of interest and the threshold corresponding to each region of interest. For example, if the threshold for the primary region of interest is set to Tz and the threshold for the secondary regions of interest is set to Tf, and only the ratio corresponding to the primary region of interest is greater than or equal to Tz, and the ratios corresponding to each of the other secondary regions of interest are all less than Tf, then it is determined that no flood has occurred. If N1 / Nz corresponding to the primary region of interest is greater than or equal to Tz, and the ratio corresponding to at least one of the other secondary regions of interest (e.g., all secondary regions of interest) is greater than or equal to Tf, then it is determined that a flood has occurred. The severity of the flood may also be determined based on the number of secondary regions of interest whose ratios are greater than or equal to Tf. For example, if there are M secondary regions of interest (where M is an integer greater than or equal to 2), and the ratios corresponding to all secondary regions of interest are greater than or equal to Tf, then it is determined that a serious flood has occurred. If the ratio corresponding to only one secondary region of interest is greater than or equal to Tf, then it is determined that a mild flood has occurred. If the ratios corresponding to secondary regions of interest that are greater than 1 and less than M are greater than or equal to Tf, then it is determined that a moderate flood has occurred. The embodiments of the present invention are not limited thereto, and the severity of the flood may be divided into more than three levels, with the specific determination method being the same, and its explanation is omitted here. As shown in Figure 9, if the ratios corresponding to all areas of interest are greater than the threshold, it is determined that a flood has occurred. As shown in Figure 8, if only the ratio corresponding to the primary area of ​​interest is greater than the threshold, and the ratios corresponding to other secondary areas of interest are less than the threshold, it is determined that a flood has not occurred.

[0059] This allows for improved accuracy in flood detection by comparing the ratio of the water area areas corresponding to multiple regions of interest with the respective thresholds for each region of interest to determine whether or not a flood has occurred.

[0060] The following explains how to determine whether or not standing water has occurred.

[0061] In some embodiments, unlike flood detection, when detecting standing water on a road, the primary region of interest includes at least the road area. Preferably, the primary region of interest may further include the sidewalk area. The secondary region of interest is optional and includes other high terrain areas. Each of the above regions of interest is pre-configured before standing water detection. Figure 10A is a schematic diagram of the pre-configured regions of interest in the original image. As shown in Figure 10A, the image data includes one primary region of interest (primary ROI) and one secondary region of interest (secondary ROI). The primary region of interest includes the road area and the sidewalk area, and the secondary region of interest includes the slope vegetation area.

[0062] In some embodiments, unlike flood detection methods, when detecting standing water on a road, it may be determined whether or not standing water has occurred (on the road) based on the ratio of the number of pixels of water to the number of pixels of a predetermined region in a given region of interest (i.e., one region of interest) in the division result of that region of interest.

[0063] In some embodiments, the calculation unit 102 may determine the ratio of the number of pixels of water to the number of pixels of a predetermined region in the primary region of interest, based on the water mask, road mask, and primary region of interest described above. The predetermined region refers to the road region (hereinafter referred to as the normal road mask mask3) separated from the normal scenario (i.e., the scenario in which no standing water occurs). For example, let N1 be the number of pixels with a pixel value of 1 in mask1 within the primary region of interest, and let Nz be the number of pixels in mask3. The determination unit 103 determines whether or not standing water has occurred based on the result of comparing the ratio of water in the primary region of interest with a threshold corresponding to the primary region of interest. For example, if the threshold corresponding to the primary region of interest is set to Tz, and the ratio N1 / Nz corresponding to the primary region of interest is greater than or equal to Tz, it is determined that standing water has occurred on the road; otherwise, it is determined that no standing water has occurred on the road. Figure 10B is another schematic diagram of the pre-set region of interest of the original image. As shown in Figure 10B, if the image data includes one primary region of interest (primary ROI) and the ratio corresponding to that primary region of interest is greater than a threshold, it is determined that water has accumulated on the road.

[0064] In some embodiments, whether or not water accumulation has occurred may be determined by comparing the ratio of the area of ​​the water region corresponding to multiple regions of interest with the respective threshold values ​​for each of the multiple regions of interest. The specific embodiment is similar to that of the flood detection method; that is, the calculation unit 102 may determine, based on the water mask, road mask and each region of interest, the ratio of the number of pixels of water to the number of pixels of road in each region of interest (which may be considered as the ratio of the area of ​​the water region to the area of ​​road). For example, let N1 be the number of pixels with a pixel value of 1 in mask1 within the main region of interest, and Nz be the number of pixels with a pixel value of 1 in mask3. Let N2 be the number of pixels with a pixel value of 1 in mask1 within one auxiliary region of interest (e.g., the slope vegetation region in Figure 6), and Nf be the number of pixels with a pixel value of 1 in mask3. For the main region of interest, the above ratio is N1 / Nz, and for the auxiliary region of interest, the above ratio is N2 / Nf. If there are other auxiliary regions of interest, the same method is used to determine the ratio for each auxiliary region of interest, and the explanation is omitted here. The determination unit 103 determines whether or not standing water has occurred based on the comparison result between the ratio of water in each area of ​​interest and the threshold corresponding to each area of ​​interest. For example, if the threshold corresponding to the primary area of ​​interest is set to Tz and the threshold corresponding to the secondary area of ​​interest is set to Tf, and the ratio corresponding to the primary area of ​​interest is smaller than Tz, and the ratios corresponding to each of the other secondary areas of interest are all smaller than Tf, then it is determined that no standing water has occurred. If N1 / Nz corresponding to the primary area of ​​interest is greater than or equal to Tz, and the ratio corresponding to at least one of the other secondary areas of interest (e.g., all secondary areas of interest) is greater than or equal to Tf, then it is determined that both the road and the semantic classes corresponding to the secondary areas of interest are standing water. If N1 / Nz corresponding to the primary area of ​​interest is greater than or equal to Tz, and the ratios corresponding to the other secondary areas of interest are all smaller than Tf, then it is determined that standing water has occurred only on the road, and not in the other areas.

[0065] In the above embodiment, each region of interest may also be represented as a mask (hereinafter abbreviated as ROI mask mask4). The size of ROI mask mask4 is the same as the size of each frame image data. In mask4, the pixel values ​​of the region of interest are set to 1, and the pixel values ​​of other positions are set to 0. Therefore, when determining the number of pixels with a pixel value of 1 in mask1 and the number of pixels in mask3 within the main region of interest, we may multiply mask4 and mask1 and then define the number of pixels with a multiplication result of 1 as N1, and multiply mask4 and mask3 and then define the number of pixels with a multiplication result of 1 as Nz, but we will omit the explanation here.

[0066] In the present invention, each of the above thresholds may be determined according to actual demand, and the embodiments of the present invention are not limited thereto.

[0067] In some embodiments, the device may further include an alarm unit (not shown, optional). The alarm unit issues an alarm if a flood or standing water is detected. The alarm may be audible, visual, or text-based, but embodiments of the present invention are not limited thereto.

[0068] The above describes only the components or modules related to the present invention, but the present invention is not limited thereto. The flood or stagnant water detection device 100 according to an embodiment of the present invention may further include other components or modules. For specific details of these components or modules, refer to related technologies.

[0069] Furthermore, for simplicity, Figures 1 and 4 only illustrate the connection relationships or signal directions between each component or module; however, those skilled in the art may employ various related techniques, such as bus connections. Each of the above components or modules may be implemented by hardware devices such as processors, memory, transmitters, and receivers, but the embodiments of the present invention are not limited thereto.

[0070] According to this embodiment, by considering contextual information and / or correlations between each positional pixel during semantic segmentation, more important information can be learned, better feature vectors for semantic segmentation can be extracted, the negative impact of imbalanced training datasets (imbalance in the number of samples in each class) on semantic segmentation can be mitigated, classification accuracy can be improved, and the accuracy of flood or stagnant water detection can be improved.

[0071] <Example 2> Embodiments of the present invention further provide electronic devices, and Figure 11 is a schematic diagram of one such electronic device according to an embodiment of the present invention. As shown in Figure 11, the electronic device 1100 includes a flood or standing water detection device 1101, the configuration and function of which are the same as those described in Embodiment 1, and are therefore omitted from this description.

[0072] In this embodiment, the electronic device 1100 may be various types of electronic devices, such as an in-vehicle terminal, a mobile terminal, or a computer.

[0073] Figure 12 is a schematic block diagram of one system configuration of an electronic device according to Embodiment 2 of the present invention. As shown in Figure 12, the electronic device 1200 may include a processor 1201 and a memory 1202, the memory 1202 being connected to the processor 1201. The figure is merely illustrative, and the configuration may be supplemented or replaced with other types of configurations to implement telecommunications functions or other functions.

[0074] As shown in Figure 12, the electronic device 1200 may further include an input unit 1203, a display 1204, and a power supply 1205.

[0075] In one embodiment of the present invention, the functions of the flood or standing water detection device of Embodiment 1 may be integrated into a processor 1201. Here, the processor 1201 may be configured to divide each frame image data of a video stream into different semantic classes based on the correlation between each pixel in each frame image data of the video stream and all position pixels, and / or contextual information in each frame image data of the video stream, calculate the percentage of water occupancy in at least one region of interest based on the division results, and detect whether or not a flood or standing water has occurred based on the percentage of water occupancy in at least one region of interest.

[0076] In some embodiments, the configuration of the processor 1201 may refer to Embodiment 1, and its description is omitted here.

[0077] In another embodiment of the present invention, the flood or standing water detection device 100 of Embodiment 1 may be arranged with a processor 1201, for example, the flood or standing water detection device 100 may be a chip connected to the processor 1201 and configured to realize the functions of the flood or standing water detection device 100 by control of the processor 1201.

[0078] In one embodiment of the present invention, the electronic device 1200 does not have to include all the components shown in Figure 12.

[0079] As shown in Figure 12, the processor 1201 is also referred to as a controller or operation control unit, and may include a microprocessor or other processing unit and / or logic unit, and the processor 1201 receives input and controls the operation of each part of the electronic device 1200.

[0080] Memory 1202 may be, for example, a buffer, flash memory, a hard disk, a portable medium, volatile memory, non-volatile memory, or one or more of other suitable devices. The processor 1201 may execute a program stored in memory 1202 to store or process information. Other components are similar to those in the prior art and are therefore omitted from this description. Each part of the electronic device 1200 may be implemented by dedicated hardware, firmware, software, or a combination thereof, without departing from the scope of the present invention.

[0081] According to this embodiment, by considering contextual information and / or correlations between each positional pixel during semantic segmentation, more important information can be learned, better feature vectors for semantic segmentation can be extracted, the negative impact of imbalanced training datasets on semantic segmentation can be mitigated, classification accuracy can be improved, and the accuracy of flood or stagnant water detection can be improved.

[0082] <Example 3> Embodiments of the present invention further provide a method for detecting flood or stagnant water corresponding to the flood or stagnant water detection device of Embodiment 1. Figure 13 is a schematic diagram of one method for detecting flood or stagnant water according to an embodiment of the present invention. As shown in Figure 13, the method includes the following steps.

[0083] Step 1301: Based on the correlation between each pixel and all position pixels in each frame image data of the video stream, and / or the context information in each frame image data of the video stream, divide each frame image data of the video stream into different semantic classes.

[0084] Step 1302: Calculate the water occupancy rate in at least one region of interest based on the partitioning results.

[0085] Step 1303: Detect whether a flood or stagnant water has occurred based on the water occupancy rate in at least one region of interest.

[0086] In some embodiments, the aspects of steps 1301 to 1303 may refer to the division unit 101, calculation unit 102, and determination unit 103 according to Embodiment 1, and the explanation of overlapping parts will be omitted.

[0087] Figure 14 is a schematic diagram of one embodiment of step 1301. As shown in Figure 14, step 1301 includes the following steps:

[0088] Step 1401: Input each frame image data into the feature extraction network and extract a feature map.

[0089] Step 1402: Determine the correlation between any two locations on the feature map.

[0090] Step 1403: Determine weights based on the correlation and update the features at each location on the feature map based on the weights.

[0091] Step 1404: Contextual information is extracted from the updated features, and the contextual information and the updated features are input into the prediction network to obtain the result of semantic class division for each frame image data.

[0092] In some embodiments, the aspects of steps 1401 to 1404 may refer to the extraction module 201, determination module 202, update module 203, and splitting module 204 according to Embodiment 1, and the explanation of the overlapping parts will be omitted.

[0093] In some embodiments, the method may further include the following steps prior to step 1302.

[0094] The semantic class with the highest frequency of occurrence among the semantic classes at the same position in the first number of divided frame image data is counted, and this semantic class with the highest frequency of occurrence is set as the semantic class for that position.

[0095] In some embodiments, the method may further include the following steps prior to step 1302.

[0096] Noise reduction processing is applied to the splitting results.

[0097] In some embodiments, the method may further include the following steps prior to step 1302.

[0098] Based on the second set of frame image data, moving objects in the image are determined, and these moving objects are removed from the division result.

[0099] In some embodiments, the method may further include the following steps prior to step 1302.

[0100] When flooding is detected, the water mask and riverbed mask are determined based on the division results, and when stagnant water is detected, the water mask and road mask are determined based on the division results.

[0101] In some embodiments, if a flood is detected in step 1303, the ratio of the number of pixels of water in each of at least two regions of interest to the number of pixels in that region of interest is calculated, and it is determined whether or not a flood has occurred based on the result of comparing this ratio of water in each region of interest with a threshold corresponding to each region of interest.

[0102] In some embodiments, when detecting standing water in step 1303, the ratio of the number of pixels in the primary region of interest to the number of pixels in a predetermined region within the primary region of interest is calculated, and based on the comparison of this ratio with a threshold, it is determined whether or not standing water has occurred.

[0103] For example, if the at least two regions of interest include at least one primary region of interest, and if flooding is detected, the primary region of interest includes at least a river region; if stagnant water is detected, the primary region of interest includes at least a road region.

[0104] Figure 15 is another schematic diagram of a flood or stagnant water detection method according to this embodiment. As shown in Figure 15, the method includes the following steps.

[0105] Step 1501: Obtain image data for each frame of the video stream and perform preprocessing on each image data frame.

[0106] Step 1502: Based on the correlation between each pixel and all position pixels in the image data of each frame of the video stream, and / or the contextual information in the image data of each frame of the video stream, divide the image data of each frame of the video stream into different semantic classes.

[0107] Step 1503 (Optional): Perform color mapping on the image data after semantic segmentation to generate a pseudo-image for display.

[0108] Step 1504: Among the semantic classes at the same position in the divided first number of frame image data, the semantic class with the highest frequency of occurrence is counted as the semantic class for that position.

[0109] Step 1505: Post-processing is performed on the segmentation results. The post-processing includes noise reduction, determining moving objects in the image based on a second number of frame image data and removing the moving objects from the segmentation results, and / or, if floods are detected, determining water masks and riverbed masks based on the segmentation results, and if stagnant water is detected, determining water masks and road masks based on the segmentation results.

[0110] Step 1506: Calculate the water occupancy rate in at least one region of interest based on the partitioning results.

[0111] Step 1507: Detect whether a flood or stagnant water has occurred based on the water occupancy rate in at least one region of interest.

[0112] Step 1508: If flooding or standing water is detected, an alarm is issued. The alarm may be audible, visual, or text-based, and embodiments of the present invention are not limited to these.

[0113] It should be noted that Figures 13 to 15 merely provide a schematic illustration of embodiments of the present invention, and the present invention is not limited thereto. For example, the execution order of each step may be appropriately adjusted, other steps may be added, or some steps may be deleted. Those skilled in the art may make appropriate modifications in accordance with the above description, and the invention is not limited to the description in Figures 13 to 15.

[0114] According to this embodiment, by considering contextual information and / or correlations between each positional pixel during semantic segmentation, more important information can be learned, better feature vectors for semantic segmentation can be extracted, the negative impact of imbalanced training datasets on semantic segmentation can be mitigated, classification accuracy can be improved, and the accuracy of flood or stagnant water detection can be improved.

[0115] Embodiments of the present invention further provide a computer-readable program that, when executed in a flood or standing water detection device or electronic device, causes a computer to execute the flood or standing water detection method described in Embodiment 3 above in the said flood or standing water detection device or electronic device.

[0116] Embodiments of the present invention further provide a storage medium for storing a computer-readable program in which a computer causes a flood or standing water detection device or electronic device to execute the flood or standing water detection method described in Embodiment 3 above.

[0117] The flood or standing water detection method implemented in the flood or standing water detection device or electronic device described with reference to embodiments of the present invention may be implemented in hardware, software modules executed by a processor, or a combination of both. For example, one or more of the functional block diagrams shown in Figure 1, or one or more combinations of the functional block diagrams, may correspond to each software module in a computer program flow, or to each hardware module. These software modules may correspond to each of the steps shown in Figure 13. These hardware modules may be implemented by hardwareizing these software modules, for example, using a field-programmable gate array (FPGA).

[0118] The software module may reside in RAM memory, flash memory, ROM memory, EPROM memory, EEPROM memory, registers, hard disk, mobile hard disk, CD-ROM, or any other form of storage medium known to those skilled in the art. The storage medium may be connected to the processor so that the processor can read information from or write information to the storage medium, or the storage medium may be a component of the processor. The processor and the storage medium reside in an ASIC. The software module may be stored in the memory of the mobile terminal or on a memory card inserted into the mobile terminal. For example, if the device (e.g., a mobile terminal) uses a relatively large capacity MEGA-SIM card or a high-capacity flash memory device, the software module may be stored on the MEGA-SIM card or high-capacity flash memory device.

[0119] One or more functional blocks and / or one or more combinations of functional blocks shown in Figure 1 may be implemented by a general-purpose processor, digital signal processor (DSP), application-specific integrated circuit (ASIC), field-programmable gate array (FPGA) or other programmable logic device, discrete gate or transistor logic unit, discrete hardware component, or any suitable combination thereof for performing the functions described herein. One or more functional blocks and / or one or more combinations of functional blocks shown in Figure 1 may be implemented, for example, by a combination of computing equipment, such as a combination of a DSP and a microprocessor, a combination of multiple microprocessors, one or more microprocessors combined with DSP communication, or any other configuration.

[0120] The above description has been made with reference to specific embodiments of the present invention, but the above description is merely illustrative and does not limit the scope of protection of the present invention. Various modifications and changes can be made to the present invention as long as they do not deviate from the spirit and principles of the present invention, and these modifications and changes also fall within the scope of the present invention.

[0121] Furthermore, the following additional information is disclosed regarding embodiments including the above-described examples. (Note 1) A device for detecting floods or stagnant water, A division unit that divides the image data of each frame of a video stream into different semantic classes based on the correlation between each pixel and all position pixels in the image data of each frame of the video stream, and / or the context information in the image data of each frame of the video stream. A calculation unit that calculates the water occupancy rate in at least one region of interest based on the division results, An apparatus comprising: a determination unit that detects whether a flood or stagnant water has occurred based on the water occupancy rate in at least one region of interest. (Note 2) The aforementioned divided portion is Each frame image data is input into the feature extraction network, and a feature map is extracted. Determine the correlation between any two locations on the feature map. Based on the aforementioned correlation, weights are determined, and based on the aforementioned weights, the features at each position in the feature map are updated. The apparatus described in Appendix 1, which extracts context information from updated features, inputs the context information and updated features into a prediction network, and obtains the result of dividing the semantic class of each frame image data. (Note 3) The apparatus according to Appendix 1, further comprising: a first processing unit that counts the semantic class with the highest frequency of occurrence among the semantic classes of the same position in a first number of frame image data divided by the division unit, and sets the semantic class with the highest frequency of occurrence as the semantic class of the position. (Note 4) The apparatus as described in Appendix 1, further comprising a second processing unit for performing noise reduction processing on the division result. (Note 5) The apparatus according to Appendix 1, further comprising a background extraction unit that determines moving objects in an image based on a second number of frame image data and removes the moving objects from the division result. (Note 6) The apparatus as described in Appendix 1, further comprising a third processing unit that, when a flood is detected, determines a water mask and a riverbed mask based on the division results, and when stagnant water is detected, determines a water mask and a road mask based on the division results. (Note 7) The calculation unit calculates the ratio of the number of pixels in each of the at least two regions of interest to the number of pixels in the region of interest, where the division result in the region of interest is water. The device described in Appendix 1, wherein the determination unit determines whether or not a flood has occurred based on the result of comparing the ratio of water in each region of interest with a threshold corresponding to each region of interest. (Note 8) The calculation unit calculates the ratio of the number of pixels in the main region of interest to the number of pixels in a predetermined region within the main region of interest, or the ratio of the number of pixels in the division result in the main region of interest to the number of pixels in the main region of interest. The apparatus as described in Appendix 1, wherein the determination unit determines whether or not accumulated water has occurred based on the result of comparing the ratio and the threshold. (Note 9) When detecting a flood, the at least two regions of interest include at least one primary region of interest, and the primary region of interest includes at least a river region. When detecting stagnant water, the device according to Appendix 7 or 8, wherein the primary area of ​​interest includes at least the road area. (Note 10) A method for detecting floodwaters or stagnant water, A step of dividing each frame image data of a video stream into different semantic classes based on the correlation between each pixel and all position pixels in each frame image data of the video stream, and / or the context information in each frame image data of the video stream. A step of calculating the water occupancy rate in at least one region of interest based on the partitioning results, A method comprising the step of detecting whether a flood or stagnant water has occurred based on the percentage of water in at least one region of interest. (Note 11) The step of separating each frame image data of a video stream into different semantic classes is, The steps include inputting each frame image data into a feature extraction network and extracting a feature map, The steps include determining the correlation between any two locations on the feature map, The steps include determining weights based on the aforementioned correlation and updating the features at each location of the feature map based on the aforementioned weights, The method according to Appendix 10, comprising the steps of: extracting context information from updated features; inputting the context information and the updated features into a prediction network; and obtaining the result of semantic class division for each frame image data. (Note 12) The method according to Appendix 10, further comprising the step of counting the semantic class with the highest frequency of occurrence among the semantic classes of the same position in the divided first number of frame image data, and setting the semantic class with the highest frequency of occurrence as the semantic class of the position. (Note 13) The method according to Appendix 10, further comprising the step of performing noise reduction processing on the division result. (Note 14) The method according to Appendix 10, further comprising the steps of determining moving objects in an image based on a second number of frame image data and removing the moving objects from the division result. (Note 15) The method according to Appendix 10, further comprising the steps of determining a water mask and a riverbed mask based on the division results when a flood is detected, and determining a water mask and a road mask based on the division results when stagnant water is detected. (Note 16) The step of calculating the water occupancy rate in at least one region of interest based on the partitioning results is: The process includes the step of calculating the ratio of the number of pixels in each of at least two regions of interest to the number of pixels in the region of interest, where the division result in the region of interest is the number of pixels in water. The step of detecting whether a flood or stagnant water has occurred based on the water occupancy rate in at least one region of interest is: The method according to Appendix 10, comprising the step of determining whether or not a flood has occurred based on the result of comparing the ratio of water in each region of interest with a threshold corresponding to each region of interest. (Note 17) The step of calculating the water occupancy rate in at least one region of interest based on the partitioning results is: The step includes calculating the ratio of the number of pixels in the primary region of interest to the number of pixels in a predetermined region within the primary region of interest, or the ratio of the number of pixels in the division result in the primary region of interest to the number of pixels in the division result in water, The step of detecting whether a flood or stagnant water has occurred based on the water occupancy rate in at least one region of interest is: The method according to Appendix 10, comprising the step of determining whether or not accumulated water has occurred based on the result of comparing the ratio with a threshold. (Note 18) When detecting a flood, the at least two regions of interest include at least one primary region of interest, and the primary region of interest includes at least a river region. When detecting standing water, the method according to Appendix 16 or 17, wherein the primary area of ​​interest includes at least the road area. (Note 19) Electronic equipment including the devices described in Appendix 1.

Claims

1. A device for detecting floods or stagnant water, A division unit that divides the image data of each frame of a video stream into different semantic classes based on the correlation between each pixel and all position pixels in the image data of each frame of the video stream, and / or the context information in the image data of each frame of the video stream. A calculation unit that calculates the water occupancy rate in at least one region of interest based on the division results, A determination unit that detects whether a flood or stagnant water has occurred based on the water occupancy rate in at least one region of interest, A device comprising: a first processing unit that counts the semantic class with the highest frequency of occurrence among the semantic classes of the same position in a first number of frame image data divided by the division unit, and sets the semantic class with the highest frequency of occurrence as the semantic class of the position.

2. The aforementioned divided portion is Each frame image data is input into the feature extraction network, and a feature map is extracted. Determine the correlation between any two locations on the feature map. Based on the aforementioned correlation, weights are determined, and based on the aforementioned weights, the features at each position in the feature map are updated. The apparatus according to claim 1, which extracts context information from updated features, inputs the context information and the updated features into a prediction network, and obtains the result of dividing the semantic class of each frame image data.

3. The apparatus according to claim 1, further comprising a second processing unit that performs noise reduction processing on the division result.

4. The apparatus according to claim 1, further comprising a background extraction unit that determines moving objects in an image based on a second number of frame image data and removes the moving objects from the division result.

5. A device for detecting floods or stagnant water, A division unit that divides the image data of each frame of a video stream into different semantic classes based on the correlation between each pixel and all position pixels in the image data of each frame of the video stream, and / or the context information in the image data of each frame of the video stream. A calculation unit that calculates the water occupancy rate in at least one region of interest based on the division results, A determination unit that detects whether a flood or stagnant water has occurred based on the water occupancy rate in at least one region of interest, An apparatus comprising: a third processing unit that, when a flood is detected, determines a water mask and a riverbed mask based on the division results; and when stagnant water is detected, determines a water mask and a road mask based on the division results.

6. A device for detecting floods or stagnant water, A division unit that divides the image data of each frame of a video stream into different semantic classes based on the correlation between each pixel and all position pixels in the image data of each frame of the video stream, and / or the context information in the image data of each frame of the video stream. A calculation unit that calculates the water occupancy rate in at least one region of interest based on the division results, Includes a determination unit that detects whether a flood or stagnant water has occurred based on the water occupancy rate in at least one region of interest, The calculation unit calculates the ratio of the number of pixels in each of the at least two regions of interest to the number of pixels in the region of interest, where the division result in the region of interest is water. The determination unit is a device that determines whether or not a flood has occurred based on the result of comparing the ratio of water in each region of interest with a threshold corresponding to each region of interest.

7. A device for detecting floods or stagnant water, A division unit that divides the image data of each frame of a video stream into different semantic classes based on the correlation between each pixel and all position pixels in the image data of each frame of the video stream, and / or the context information in the image data of each frame of the video stream. A calculation unit that calculates the water occupancy rate in at least one region of interest based on the division results, Includes a determination unit that detects whether a flood or stagnant water has occurred based on the water occupancy rate in at least one region of interest, The calculation unit calculates the ratio of the number of pixels in the main region of interest to the number of pixels in a predetermined region within the main region of interest, or the ratio of the number of pixels in the division result in the main region of interest to the number of pixels in the main region of interest. The determination unit is a device that determines whether or not water has accumulated based on the result of comparing the ratio with a threshold.

8. When detecting a flood, the at least two regions of interest include at least one primary region of interest, and the primary region of interest includes at least a river region. The apparatus according to claim 6, wherein, when detecting stagnant water, the primary area of ​​interest includes at least the road area.

9. A method for detecting floodwaters or stagnant water, A step of dividing each frame image data of a video stream into different semantic classes based on the correlation between each pixel and all position pixels in each frame image data of the video stream, and / or the context information in each frame image data of the video stream. A step of calculating the water occupancy rate in at least one region of interest based on the partitioning results, A step of detecting whether a flood or stagnant water has occurred based on the water occupancy rate in at least one region of interest, A method comprising the steps of counting the semantic class with the highest frequency of occurrence among the semantic classes of the same position in a first number of divided frame image data, and setting the semantic class with the highest frequency of occurrence as the semantic class of the position.