A water accumulation condition detection method, device, equipment and storage medium
By identifying water accumulation areas in video frames and combining them with depth maps, the depth of water accumulation can be automatically measured, solving the problems of high measurement cost and low efficiency in existing technologies, and realizing water depth measurement with wide applicability.
Patent Information
- Application Number
- CN202211418483.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-11-14
- Publication Date
- 2026-02-06
- Estimated Expiration
- 2042-11-14
AI Technical Summary
Existing methods for measuring water depth are costly, inefficient, and incompatible with waterlogged areas without water gauges or reference points.
By acquiring target video frames from video segments, identifying waterlogged areas and combining them with depth maps for precise location, the system automatically measures the depth of the water without requiring manual measurement equipment or reference objects.
It enables automatic and accurate location and depth measurement of waterlogged areas, reduces measurement costs, improves efficiency, and is suitable for various waterlogged areas.
Smart Images

Figure CN115661721B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of water accumulation detection, and in particular to a water accumulation detection method, device, equipment and storage medium. BACKGROUND
[0002] The waterlogging phenomenon in the city has always been one of the problems that must be solved in the process of urban development. When there is heavy rainfall for a short time or extreme weather such as continuous rainfall, the rainfall exceeds the drainage capacity of the city, and the road waterlogging phenomenon may occur in the city. This phenomenon will limit the function of the city's transportation road and other infrastructure, and in severe cases, it will even cause traffic paralysis, which will have a negative impact on people's economic life.
[0003] In view of the problem of waterlogging in urban roads, in addition to increasing the drainage capacity of the city and taking preventive measures before water accumulation occurs, timely warning of road waterlogging can also be made when extreme weather such as heavy rainfall occurs. To a certain extent, the warning of road waterlogging can eliminate or reduce the hidden dangers caused by road waterlogging.
[0004] It can be understood that in order to realize the warning of road waterlogging, it is necessary to measure the water depth. At present, the main way to measure the water depth is the manual measurement method, that is, the measurement personnel hold the measurement equipment to the water accumulation area to measure the water depth. However, the manual measurement method has high measurement cost (including labor cost and equipment cost) and low measurement efficiency. SUMMARY
[0005] Therefore, the present application provides a water accumulation detection method, device, equipment and storage medium to solve the problem of high measurement cost and low measurement efficiency of the existing water depth measurement method. The technical solution is as follows:
[0006] A water accumulation detection method, comprising:
[0007] Obtaining a target video segment under a specified video point;
[0008] Identifying the water accumulation area in each video frame contained in the target video segment as a preliminary water accumulation area;
[0009] Obtaining the depth map corresponding to each video frame contained in the target video frame set, wherein the video frame contained in the target video frame set is the video frame in which the water accumulation area is identified;
[0010] According to the preliminary water accumulation area in each video frame contained in the target video frame set and the depth map corresponding to each video frame in the target video frame set, the target water accumulation area in the target video frame is located, wherein the target video frame is a video frame in the target video frame set;
[0011] determining pixel values of a region corresponding to the target waterlogging region in a depth map corresponding to the target video frame as depth information of the target waterlogging region.
[0012] Optionally, the target waterlogging region in the target video frame is located according to the preliminary waterlogging region in each video frame included in the target video frame set and the depth map corresponding to each video frame in the target video frame set, and the method comprises the following steps.
[0013] For each video frame included in the target video frame set: correcting the preliminary waterlogging region in the video frame according to a reference depth map and a depth map corresponding to the video frame, to obtain a corrected waterlogging region in the video frame, wherein the reference depth map is determined based on a reference video frame set, and the reference video frame set comprises video frames without waterlogging regions obtained from the video under the specified video point.
[0014] determining an intersection region and a union region of the corrected waterlogging regions in the video frames included in the target video frame set;
[0015] locating the target waterlogging region in the target video frame according to the intersection region, the union region and the depth map corresponding to the target video frame.
[0016] Optionally, the step of correcting the preliminary waterlogging region in the video frame according to the reference depth map and the depth map corresponding to the video frame comprises the following steps.
[0017] For each pixel point included in the preliminary waterlogging region in the video frame:
[0018] obtaining a depth value corresponding to the pixel point from the depth map corresponding to the video frame, and obtaining a reference depth value corresponding to the pixel point from the reference depth map;
[0019] calculating a difference value between the depth value corresponding to the pixel point and the reference depth value corresponding to the pixel point as a depth difference value corresponding to the pixel point;
[0020] if the depth difference value corresponding to the pixel point is less than a preset waterlogging depth threshold, determining that the pixel point does not belong to a waterlogging region, and removing the pixel point from the preliminary waterlogging region.
[0021] Optionally, the process of determining the reference depth map based on the reference video frame set comprises the following steps.
[0022] obtaining a depth map corresponding to each video frame included in the reference video frame set;
[0023] Average the pixel values of the same positions in the depth maps corresponding to the video frames included in the reference video frame set, to obtain an average depth map, and the average depth map is taken as the reference depth map.
[0024] Optionally, the target water area in the target video frame is located according to the intersection region, the union region, and the depth map corresponding to the target video frame, including:
[0025] Obtain the depth values corresponding to the respective pixels in the first water area and the second water area in the target video frame from the depth map corresponding to the target video frame, wherein the first water area and the second water area are the regions corresponding to the intersection region and the union region in the target video frame, respectively;
[0026] Average the depth values corresponding to the respective pixels in the first water area to obtain an average depth value;
[0027] Calculate the difference between the depth values corresponding to the respective pixels in the second water area and the average depth value to obtain the depth difference values corresponding to the respective pixels in the second water area;
[0028] Determine the pixels in the second water area as non-target pixels if the corresponding depth values of the pixels are greater than a preset depth change threshold;
[0029] Determine the region composed of the pixels in the second water area except the non-target pixels as the target water area in the target video frame.
[0030] Optionally, the water area in each video frame included in the target video segment is identified, including:
[0031] For each video frame included in the target video segment:
[0032] Predict the respective categories to which the respective pixels included in the video frame belong, wherein the category to which a pixel belongs is one of water and background;
[0033] Determine the region composed of the pixels in the video frame whose categories are water as the water area identified from the video frame.
[0034] Optionally, the target video segment is captured based on a monocular camera.
[0035] The depth map corresponding to each video frame included in the target video frame set is obtained, including:
[0036] For each video frame in the target video frame set:
[0037] generate a disparity change map corresponding to the video frame based on the video frame, wherein the disparity change map corresponding to the video frame is a right image generated for a left image being the video frame;
[0038] generate a disparity map corresponding to the video frame based on the video frame and the disparity change map corresponding to the video frame;
[0039] determine a depth map corresponding to the video frame based on the disparity map corresponding to the video frame.
[0040] Optionally, the generating the disparity change map corresponding to the video frame based on the video frame comprises:
[0041] extracting a plurality of hierarchical multi-channel feature maps from the video frame, and processing the plurality of hierarchical multi-channel feature maps into a plurality of target feature maps having the same size as the video frame, wherein each target feature map contains feature information of the same channel at different levels;
[0042] predicting a plurality of disparity probability maps under a plurality of preset disparity offset values based on the plurality of target feature maps, and respectively offsetting positions of each pixel value of the video frame based on the plurality of disparity offset values to obtain a plurality of disparity offset maps under the plurality of disparity offset values;
[0043] generating the disparity change map corresponding to the video frame based on the plurality of disparity probability maps under the plurality of disparity offset values and the plurality of disparity offset maps under the plurality of disparity offset values.
[0044] Optionally, the generating the disparity change map corresponding to the video frame based on the plurality of disparity probability maps under the plurality of disparity offset values and the plurality of disparity offset maps under the plurality of disparity offset values comprises:
[0045] multiplying the disparity probability map and the disparity offset map under the same disparity offset value to obtain a plurality of multiplication results;
[0046] fusing the plurality of multiplication results, and taking a fusion result as the disparity change map corresponding to the video frame.
[0047] Optionally, the generating the disparity change map corresponding to the video frame based on the video frame comprises:
[0048] processing the video frame based on a pre-trained disparity change map generation model to generate the disparity change map corresponding to the video frame;
[0049] wherein the disparity change map generation model is trained based on a training video frame as a training sample and a real disparity change map corresponding to the training video frame as a sample label, the training video frame is a left image collected based on a binocular camera, and the real disparity change map corresponding to the training video frame is a right image collected based on the binocular camera.
[0050] Optionally, the generating the disparity map corresponding to the video frame based on the video frame and the disparity change map corresponding to the video frame comprises:
[0051] processing the video frame and the disparity change map corresponding to the video frame based on a pre-trained disparity map generation model to generate the disparity map corresponding to the video frame.
[0052] The disparity map generation model is trained using training video frames and training disparity change maps corresponding to the training video frames as training samples, and using real disparity maps corresponding to the training video frames as sample labels.
[0053] A water accumulation condition detection device comprises a target video segment acquisition module, a water accumulation area identification module, a depth map acquisition module, a target water accumulation area determination module, and a water accumulation depth information determination module.
[0054] The target video segment acquisition module is configured to acquire a target video segment at a specified video point.
[0055] The water accumulation area identification module is configured to identify water accumulation areas in each video frame included in the target video segment as preliminary water accumulation areas.
[0056] The depth map acquisition module is configured to acquire a depth map corresponding to each video frame included in a target video frame set, wherein the video frames included in the target video frame set are video frames in which water accumulation areas are identified.
[0057] The target water accumulation area determination module is configured to locate a target water accumulation area in a target video frame based on preliminary water accumulation areas in each video frame included in the target video frame set and the depth map corresponding to each video frame in the target video frame set, wherein the target video frame is a video frame in the target video frame set.
[0058] The water accumulation depth information determination module is configured to determine a pixel value of a region corresponding to the target water accumulation area in the depth map corresponding to the target video frame as depth information of the target water accumulation area.
[0059] A water accumulation condition detection device comprises a memory and a processor.
[0060] The memory is configured to store a program.
[0061] The processor is configured to execute the program to implement each step of the water accumulation condition detection method described in any of the above embodiments.
[0062] A computer readable storage medium, having stored thereon a computer program, which, when executed by a processor, implements each step of the waterlogging condition detection method of any one of the preceding aspects.
[0063] The waterlogging condition detection method, device, equipment and storage medium provided by the application first acquire a target video segment under a specified video point, then identify waterlogging areas in each video frame contained by the target video segment, considering that the identified waterlogging areas are not accurate enough, the application takes the identified waterlogging areas as preliminary waterlogging areas, then acquires a depth map corresponding to each video frame of a target video frame set (the video frames contained by the target video frame set are the video frames in which waterlogging areas are identified), then locates a target waterlogging area in the target video frame according to the preliminary waterlogging areas in each video frame of the target video frame set and the depth map corresponding to each video frame of the target video frame set, and finally determines the pixel value of the area corresponding to the target waterlogging area in the depth map corresponding to the target video frame as the depth information of the target waterlogging area. The waterlogging condition detection method provided by the application can automatically and accurately locate the waterlogging area in the video frame contained by the target video segment, and on this basis, the depth information of the located waterlogging area can be further determined. The waterlogging condition detection method provided by the application does not require a measurement personnel to hold a measurement device to measure the waterlogging depth of the waterlogging area, and compared with the manual measurement method, the measurement cost is greatly reduced and the measurement efficiency is improved. BRIEF DESCRIPTION OF DRAWINGS
[0064] In order to more clearly illustrate the technical solutions in the embodiments of the application or the prior art, below the drawings needed to be used in the embodiments or prior art description will be briefly introduced. Obviously, the drawings in the following description are only embodiments of the application, and for those skilled in the art, other drawings can be obtained without creative labor on the basis of the provided drawings.
[0065] Figure 1 The schematic diagram of the hardware architecture related to the application;
[0066] Figure 2 The flowchart of the waterlogging condition detection method provided by the embodiment of the application;
[0067] Figure 3 An example of the waterlogging area identified from a video frame based on the semantic segmentation model provided by the embodiment of the application;
[0068] Figure 4 The flowchart of the depth map corresponding to a video frame provided by the embodiment of the application;
[0069] Figure 5 An example of the disparity change map generation model provided by the embodiment of the application;
[0070] Figure 6 A flowchart for locating a target waterlogging region in a target video frame set according to a preliminary waterlogging region in each video frame included in the target video frame set and a depth map corresponding to each video frame in the target video frame set;
[0071] Figures 7(a) to 7(c) A schematic diagram of a corrected waterlogging region in three video frames according to an embodiment of the present application;
[0072] Figures 8(a) to 8(b) A schematic diagram of an intersection region and a union region of the corrected waterlogging region in the three video frames shown in FIG. 7;
[0073] Figure 9 A schematic diagram of a waterlogging condition detection device according to an embodiment of the present application;
[0074] Figure 10 A schematic diagram of a waterlogging condition detection device according to an embodiment of the present application. DETAILED DESCRIPTION
[0075] The technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only some of the embodiments of the present application, but not all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative work fall within the scope of the present application.
[0076] Considering that the existing waterlogging depth measurement method has high measurement cost and low measurement efficiency, the present inventor attempts to propose a scheme capable of reducing the measurement cost and improving the measurement efficiency. Therefore, research is conducted, and the initial idea is as follows:
[0077] An automatic measurement scheme combining a water gauge and waterlogging segmentation is adopted. The general process of the scheme is as follows: an image including a physical water gauge is collected, water gauge recognition is performed on the image including the physical water gauge to obtain scale data of the physical water gauge, a virtual water gauge is generated based on the scale data of the physical water gauge, waterlogging region segmentation is performed on the image including the physical water gauge to obtain water level line position information, and depth information of the waterlogging region is calculated based on the virtual water gauge and the water level line position information.
[0078] The above scheme does not require manual holding of a measurement device to measure the waterlogging depth of the waterlogging region, thereby reducing the measurement cost and improving the measurement efficiency. However, the above scheme can obtain relatively accurate waterlogging depth information for the water gauge camera point position. However, for the waterlogging region under the point position of the water gauge-free device, the waterlogging depth information cannot be measured. It can be seen that the above measurement scheme combining the water gauge and the waterlogging segmentation has great limitations.
[0079] In view of the above scheme has certain defects, the present inventor continues to research, in the research process thought, can help reference object realizes the measurement of water depth, specifically, in the video frame of the video data contains the pedestrian in the water area as a reference object, through target detection and image segmentation method obtains the pedestrian information in the video frame, helps the pedestrian information obtained to determine the water depth information of the water area.
[0080] The above measurement scheme with reference object although does not need to utilize the water gauge, but also has limitation, it can only measure the water depth of the water area with reference object, cannot be compatible with the water area without reference object, essentially still does not solve the problem of measurement limitation.
[0081] Therefore, the present inventor continues to research, through continuous research, finally proposes a water condition detection method, which can detect the water area in the video frame contained in the video segment to be analyzed and the water depth information of the detected water area by analyzing the video segment, the detection method does not need to measure the water depth by measuring personnel holding measuring equipment to the water area, also does not need to help water gauge and reference object.
[0082] Before introducing the water condition detection method provided by the present application, the hardware architecture involved in the present application is described.
[0083] In a possible implementation manner, as shown in Figure 1 The hardware architecture involved in the present application can include: an electronic device 101 and a server 102.
[0084] For example, the electronic device 101 can be any kind of electronic product that can interact with the user through one or more ways such as keyboard, touchpad, touch screen, remote control, voice interaction or handwriting device, for example, mobile phone, notebook computer, tablet computer, palm computer, personal computer, wearable device, smart television, PAD, etc.
[0085] It should be noted that, Figure 1 Only one example, the type of electronic device can be various, not limited to Figure 1 The notebook computer in
[0086] For example, the server 102 can be a server, or a server cluster composed of multiple servers, or a cloud computing server center. The server 102 can include a processor, a memory and a network interface, etc.
[0087] For example, the electronic device 101 can establish connection and communication with the server 102 through a wireless communication network; for example, the electronic device 101 can establish connection and communication with the server 102 through a wired network.
[0088] The electronic device 101 acquires a video segment to be analyzed at a specified video point, and sends the acquired video segment to be analyzed to the server 102. The server detects a waterlogging area in the video segment and a waterlogging depth of the waterlogging area according to the waterlogging condition detection method provided in the present application, and sends the detection result to the electronic device 101.
[0089] In another possible implementation, the hardware architecture provided in the present application can include an electronic device. The electronic device is a device with strong data processing capability.
[0090] For example, the electronic device can be any electronic product that can interact with a user through one or more of a keyboard, a touchpad, a touch screen, a remote controller, voice interaction, a handwriting device, and the like, such as a PC, a mobile phone, a notebook computer, a tablet computer, a palm computer, a personal computer, and the like.
[0091] The electronic device can detect a waterlogging area in the video segment to be analyzed and a waterlogging depth of the waterlogging area according to the waterlogging condition detection method provided in the present application, and output the detection result.
[0092] Those skilled in the art should understand that the electronic device and the server described above are only examples, and other existing or future electronic devices or servers that can be applicable to the present application should also be included in the protection scope of the present application, and are hereby included by reference.
[0093] Next, the waterlogging condition detection method provided in the present application is described below through the following embodiments.
[0094] Referring to FIG. 1, Figure 2 , a flowchart of the waterlogging condition detection method provided in the embodiment of the present application is shown, which can include the following steps.
[0095] Step S201: Acquire a target video segment at a specified video point.
[0096] The specified video point is a video point in a region to be detected, and the present application aims to detect the waterlogging condition of a waterlogging area in the region to be detected.
[0097] Specifically, the process of acquiring the target video segment at the specified video point includes: acquiring a video segment to be analyzed (for example, a video segment of a certain period on a rainy day) from a video at the specified video point as the target video segment. Optionally, the video at the specified video point can be captured based on a monocular camera.
[0098] Step S202: Identify a waterlogging area in each video frame included in the target video segment as a preliminary waterlogging area.
[0099] The target video segment includes a plurality of video frames, and the step is to identify the accumulated water area in each video frame included in the target video segment.
[0100] Optionally, the accumulated water area in each video frame included in the target video segment can be identified by using a semantic segmentation method. Since the accumulated water area is usually irregular, the range of the accumulated water area identified by using the semantic segmentation method can be more consistent with the boundary of the accumulated water area, and the interference of the non-accumulated water part on the subsequent accumulated water depth judgment can be reduced. Therefore, the accumulated water area can be identified by using the semantic segmentation method.
[0101] Specifically, for each video frame included in the target video segment, first, the category to which each pixel point included in the video frame belongs is predicted, wherein the category to which a pixel point belongs is one of accumulated water and background, and then the region composed of the pixel points in the video frame whose category is accumulated water is determined as the accumulated water area identified from the video frame.
[0102] Optionally, when predicting the category to which each pixel point included in a video frame belongs, the category to which each pixel point included in the video frame belongs can be predicted based on a pre-trained semantic segmentation model. Specifically, the video frame is input into the pre-trained semantic segmentation model (such as a semantic segmentation model based on deeplabv3), and the semantic segmentation model outputs the prediction probability corresponding to each pixel point included in the video frame (the prediction probability corresponding to a pixel point is the probability that the category to which the pixel point belongs is accumulated water and background, respectively). The category to which each pixel point included in the video frame belongs is determined according to the prediction probability corresponding to each pixel point included in the video frame. Please refer to Figure 3 , which shows an example of the accumulated water area identified from a video frame based on a semantic segmentation model, Figure 3 The black region in the figure is the identified accumulated water area.
[0103] When identifying the accumulated water area from the video frames included in the target video segment in the above manner, since only image-level information is used, the identified accumulated water area may have errors, i.e., the identified accumulated water area is not accurate enough. In view of this, the identified accumulated water area is used as a preliminary accumulated water area, rather than a final accumulated water area.
[0104] Step S203: Obtain the depth map corresponding to each video frame included in the target video frame set.
[0105] The video frames included in the target video frame set are the video frames in which the accumulated water area is identified. It should be noted that among the video frames included in the target video segment, there may be video frames in which the accumulated water area is identified, and there may also be video frames in which the accumulated water area is not identified. The target video frame set in this step includes the video frames in which the accumulated water area is identified among the video frames included in the target video segment.
[0106] For example, the target video segment contains 120 video frames, by performing water area identification on the 120 video frames, it is found that no water area is identified in 20 video frames, and water area is identified in 100 video frames, so the target video frame set includes the 100 video frames in which water area is identified.
[0107] It should be noted that the depth map d corresponding to a video frame I contains depth information corresponding to each pixel point in the video frame I, and the depth information corresponding to a pixel point p I in the video frame I is the pixel value of the pixel point d I corresponding to the pixel point p I in the depth map d corresponding to the video frame I, for example, the depth information corresponding to the pixel point in the 5th row and the 8th column in the video frame I is the pixel value of the pixel point in the 5th row and the 8th column in the depth map d corresponding to the video frame I.
[0108] Step S204: According to the preliminary water area in each video frame included in the target video frame set and the depth map corresponding to each video frame in the target video frame set, the target water area in the target video frame is located.
[0109] Among them, the target video frame is a video frame in the target video frame set. Optionally, the target video frame can be a video frame specified by the user from the target video frame set. The target water area in the target video frame is a precise water area located in the target video frame.
[0110] It should be noted that step S202 only uses image-level information to identify water area from video frames, and step S204 further accurately locates the water area in the target video frame based on the preliminary water area and the depth information corresponding to the video frame.
[0111] Step S205: Determine the pixel value of the region corresponding to the target water area in the depth map corresponding to the target video frame as the depth information of the target water area.
[0112] After locating the target water area in the target video frame, the region corresponding to the target water area in the target video frame (the region in the depth map corresponding to the target video frame has the same position as the target water area) can be determined from the depth map corresponding to the target video frame, and the pixel value of the pixel points included in the region is determined as the depth information of the target water area.
[0113] The water accumulation condition detection method provided by the embodiment of the present application firstly acquires a target video segment under a specified video point, then identifies water accumulation regions in each video frame contained by the target video segment, considering that the identified water accumulation regions are not accurate enough, the identified water accumulation regions are taken as preliminary water accumulation regions, then depth maps respectively corresponding to each video frame contained by the target video frame set (the video frames contained by the target video frame set are the video frames in which the water accumulation regions are identified) are acquired, then the target water accumulation region in the target video frame is located according to the preliminary water accumulation regions in each video frame contained by the target video frame set and the depth maps respectively corresponding to each video frame in the target video frame set, and finally the pixel value of the region corresponding to the target water accumulation region in the depth map corresponding to the target video frame is determined as the depth information of the target water accumulation region. The water accumulation condition detection method provided by the present application can automatically and accurately locate the water accumulation region in the video frame contained by the target video segment, and the depth information of the located water accumulation region can be further determined on this basis. The water accumulation condition detection method provided by the embodiment of the present application does not need a measurement personnel to hold a measurement device to measure the water depth of the water accumulation region, compared with the manual measurement method, the measurement cost is greatly reduced and the measurement efficiency is improved, in addition, the water accumulation condition detection method provided by the embodiment of the present application does not need to use a water gauge and does not need to rely on a reference object, and has wide applicability.
[0114] In another embodiment of the present application, the specific implementation process of "step S203: acquiring the depth map respectively corresponding to each video frame contained by the target video frame set" in the above embodiment is introduced.
[0115] Since the way of acquiring the depth map respectively corresponding to each video frame contained by the target video frame set is the same, the embodiment takes one video frame I in the target video frame set as an example to introduce the implementation process of acquiring the depth map corresponding to the video frame I. The above embodiment mentions that the target video segment can be a video segment acquired from a video shot based on a monocular camera, that is, the video frames in the target video frame set can be video frames collected based on a monocular camera, and the embodiment introduces the implementation process of acquiring the depth map corresponding to the video frame I collected based on the monocular camera.
[0116] Please refer to Figure 4 , which shows a flowchart of acquiring the depth map corresponding to the video frame I, which can include:
[0117] Step S401: generating the disparity change map corresponding to the video frame I based on the video frame I.
[0118] It should be noted that when shooting based on a binocular camera, two images, i.e., a left image and a right image, are obtained at the same time, and since the video frame I is shot based on a monocular camera, the present application takes the video frame I as the left image and generates the right image based on the left image.
[0119] Optionally, the specific implementation process of generating the disparity change map corresponding to the video frame I based on the video frame I can include:
[0120] Step S4011: Extracting a plurality of levels of multi-channel feature maps from the video frame I, and processing the plurality of levels of multi-channel feature maps into a plurality of target feature maps with the same size as the video frame I.
[0121] Each target feature map contains feature information of the same channel at each level.
[0122] Specifically, the process of processing the plurality of levels of multi-channel feature maps into a plurality of target feature maps with the same size as the video frame I can include:
[0123] Step a1, processing each level of multi-channel feature map into a feature map with a unified channel number and a unified size to obtain a plurality of levels of processed multi-channel feature maps.
[0124] In view of the different channel numbers and sizes of multi-channel feature maps at different levels, in order to facilitate subsequent processing, after obtaining the plurality of levels of multi-channel feature maps, each level of multi-channel feature map is processed into a feature map with a unified channel number and a unified size, for example, each level of multi-channel feature map is processed into an N-channel h*w feature map.
[0125] Step a2, fusing the feature maps at the same channel position in the plurality of levels of processed multi-channel feature maps to obtain a plurality of fused feature maps.
[0126] Assuming that each level of processed multi-channel feature map is an N-channel h*w feature map, the feature maps at the first channel position in each level of N-channel h*w feature map are fused (for example, added),..., the feature maps at the Nth channel position in each level of N-channel h*w feature map are fused, and finally N h*w fused feature maps are obtained.
[0127] Step a3, processing the plurality of fused feature maps into feature maps with the same size as the video frame I to obtain a plurality of processed feature maps as the plurality of target feature maps.
[0128] Assuming that the size of the video frame I is H*W, and N h*w fused feature maps are obtained through step a2, step a3 processes the N h*w fused feature maps into N H*W feature maps, and the N H*W feature maps are taken as N target feature maps.
[0129] Step S4012: predicting a disparity probability map under a set of multiple disparity offset values based on the plurality of target feature maps, and respectively offsetting the positions of each pixel value of the video frame I based on the set of multiple disparity offset values to obtain a disparity offset map under the multiple disparity offset values.
[0130] wherein the plurality of disparity offset values are different from each other, the number of disparity offset values is same as the number of channels of the processed multi-channel feature map (also the number of target feature maps), and the processed multi-channel feature map is an N-channel h*w feature map, then the number of disparity offset values is N.
[0131] The N target feature maps (i.e. the N h*w feature maps) are normalized to obtain N disparity probability maps as the disparity probability maps under the plurality of disparity offset values, and it is to be noted that the sum of the disparity probabilities at the same position in the N disparity probability maps is 1. It is to be noted that each pixel value in the disparity probability map under a disparity offset value is the probability of the disparity offset value corresponding to the pixel point in the video frame I.
[0132] The disparity offset map obtained by offsetting the positions of the pixel values in the video frame based on the disparity offset value d may be expressed as:
[0133]
[0134] In this embodiment, offsetting the pixel values of the pixel points in the video frame I based on the disparity offset value d means rolling the positions of the pixel values of the pixel points in the video frame I based on the disparity offset value d. For example, when d = 1, the positions of the pixel values of the pixel points in the video frame I are rolled to the right by 1 bit, and the leftmost pixels of the offset image are cyclically supplemented with the pixel values of the rightmost pixels of the original image (i.e. the video frame I); when d = -1, the positions of the pixel values of the pixel points in the video frame I are rolled to the left by 1 bit, and the rightmost pixels of the offset image are cyclically supplemented with the pixel values of the leftmost pixels of the original image (i.e. the video frame I); it is to be noted that when d = 0, the disparity offset map corresponding to the video frame I is the video frame I.
[0135] It is to be noted that the disparity offset value d is an integer, and the value of the disparity offset value d is related to the number N of channels of the processed multi-channel feature map; specifically, if N is even, then if N is odd, then
[0136] Step S4013: generating the disparity change map corresponding to the video frame I based on the disparity probability maps under the plurality of disparity offset values and the disparity offset maps under the plurality of disparity offset values.
[0137] Optionally, the disparity probability map under the same disparity offset value and the disparity offset map are multiplied first to obtain a plurality of multiplication results, and then the plurality of multiplication results are summed to obtain the sum as the disparity change map corresponding to the video frame I, i.e.
[0138]
[0139] wherein, denotes a disparity change map corresponding to the video frame I, I d denotes a disparity offset map under a disparity offset value d, D d denotes a disparity probability map under a disparity offset value d.
[0140] Optionally, the steps S4011-S4013 can be implemented based on a disparity change map generation model. Of course, the present embodiment is not limited thereto, i.e., the present embodiment does not limit the implementation form of the steps S4011-S4013.
[0141] If the disparity change map corresponding to the video frame I is generated based on the disparity change map generation model, in implementation, the video frame I is input into the disparity change map generation model, and the disparity change map generation model processes the video frame I according to the process of the steps S4011-S4013 to generate the disparity change map corresponding to the video frame I.
[0142] Referring to Figure 5 , an example of a disparity change map generation model is shown, Figure 5 The disparity change map generation model shown in
[0143] Specifically, the process of generating the disparity change map corresponding to the video frame I based on the disparity change map generation model shown in Figure 5 is as follows:
[0144] Step b1, after the video frame I is input into the disparity change map generation model, the first feature extraction module 501 extracts multiple levels of multi-channel feature maps from the input video frame I.
[0145] Optionally, the first feature extraction module 501 can adopt a convolutional neural network, i.e., multiple levels of multi-channel feature maps are extracted from the video frame I based on the convolutional neural network.
[0146] Step b2, after obtaining multiple levels of multi-channel features, the first feature processing module 502 processes the multiple levels of multi-channel features into feature maps with uniform channel numbers and uniform sizes to obtain multiple levels of processed multi-channel feature maps.
[0147] Specifically, the first feature processing module 502 can unify the feature channel numbers and sizes by upsampling the multiple levels of multi-channel feature maps. Optionally, the first feature processing module 502 can adopt a deconvolutional layer, i.e., the multiple levels of multi-channel feature maps are upsampled based on the deconvolutional layer.
[0148] Step b3, after obtaining the multiple levels of processed multi-channel feature maps, the feature fusion module 503 fuses the multiple levels of processed multi-channel feature maps to obtain multiple fused feature maps.
[0149] When fusing the multiple levels of processed multi-channel feature maps, the feature maps at the same channel position in the multiple levels of processed multi-channel feature maps are fused (such as addition).
[0150] Step b4, after obtaining the multiple fused feature maps, the second feature processing module 504 processes the multiple fused feature maps into feature maps of the same size as the video frame I to obtain multiple processed feature maps as the multiple target feature maps.
[0151] Specifically, the second feature processing module 504 can adjust the size by upsampling the multiple fused feature maps. Optionally, the second feature processing module 504 can use a deconvolution layer, i.e., based on the deconvolution layer to upsample the multiple fused feature maps.
[0152] Step b5, after obtaining the multiple target feature maps, the disparity probability map generation module 505 normalizes the multiple processed feature maps to obtain multiple disparity probability maps as the disparity probability maps under the set multiple disparity offset values, and the disparity offset map generation module 506 offsets the pixel value position of the video frame I based on the set multiple disparity offset values to obtain the disparity offset maps under the set multiple disparity offset values.
[0153] Optionally, the disparity probability map generation module 505 can use a softmax layer, i.e., based on the softmax layer to normalize the multiple fused feature maps.
[0154] Step b6, the disparity change map generation module 507 multiplies and fuses the disparity probability map and the disparity offset map under the same disparity offset value to obtain the disparity change map corresponding to the video frame I.
[0155] It should be noted that the above disparity change map generation model takes a training video frame as a training sample, and takes a real disparity change map corresponding to the training video frame as a sample label to train, wherein the training video frame is a left image collected based on a binocular camera, and the real disparity change map corresponding to the training video frame is a right image collected based on the binocular camera.
[0156] In the training of the disparity change map generation model, first, the training video frame is input into the disparity change map generation model to obtain a disparity change map output by the disparity change map generation model as a predicted disparity change map corresponding to the training video frame, then a prediction loss of the disparity change map generation model is determined according to the predicted disparity change map corresponding to the training video frame and a real disparity change map corresponding to the training video frame, and finally the parameter of the disparity change map generation model is updated according to the prediction loss of the disparity change map generation model. The disparity change map generation model is iteratively trained in the above manner by using different training video frames until a training end condition (for example, a preset training number is reached, or the performance of the disparity change map generation model meets the requirement) is met.
[0157] The prediction loss of the disparity change map generation model can be calculated based on the following formula:
[0158]
[0159] The prediction loss of the disparity change map generation model can be calculated based on the following formula: The predicted disparity change map corresponding to the training video frame is represented as I r The real disparity change map corresponding to the training video frame is represented as I
[0160] Step S402: Based on the video frame I and the disparity change map corresponding to the video frame I, a disparity map corresponding to the video frame I is generated.
[0161] Specifically, first, features can be extracted from the video frame I and the disparity change map corresponding to the video frame I, respectively, then the features extracted from the video frame I are integrated with the features extracted from the disparity change map corresponding to the video frame I, and then the integrated features are processed and restored to the size of the video frame I, that is, the disparity map corresponding to the video frame I is obtained.
[0162] Optionally, the disparity map corresponding to the video frame I can be generated based on a model, that is, a disparity map generation model. When the disparity map corresponding to the video frame I is generated based on the disparity map generation model, the video frame I and the disparity change map corresponding to the video frame I can be input into the disparity map generation model for processing to obtain the disparity map corresponding to the video frame I output by the disparity map generation model.
[0163] Optionally, the disparity map generation model can comprise a second feature extraction module, a third feature processing module, a fourth feature processing module and a fifth feature processing module. After the video frame I and the disparity change map corresponding to the video frame I are input into the disparity map generation model, the second feature extraction module is used to extract features from the video frame I and the disparity change map corresponding to the video frame I, respectively. Optionally, the second feature extraction module can adopt two convolutional neural networks sharing the same weight, one of which is used to extract features from the video frame I, and the other of which is used to extract features from the disparity change map corresponding to the video frame I, so as to obtain the features corresponding to the video frame I and the features corresponding to the disparity change map corresponding to the video frame I. Then, the third feature processing module is used to integrate and process the features corresponding to the video frame I and the features corresponding to the disparity change map corresponding to the video frame I, so as to obtain integrated features. Next, the fourth feature processing module is used to perform multiple deconvolution and bilinear interpolation processing on the integrated features. Finally, the fifth feature processing module is used to perform up-sampling processing on the output result of the fourth feature processing module, so as to restore the size of the video frame I, and obtain the disparity map corresponding to the video frame I.
[0164] The disparity map generation model described above is trained by taking the training video frame and the disparity change map corresponding to the training video frame as training samples, and taking the real disparity map corresponding to the training video frame as a sample label. The disparity change map corresponding to the training video frame can be generated based on the disparity change map generation model described above. Optionally, the real disparity map corresponding to the training video frame can be determined based on the depth map corresponding to the training video frame. The depth map corresponding to the training video frame can be obtained based on a ranging device (such as a depth camera). After obtaining the depth map corresponding to the training video frame, the depth map corresponding to the training video frame can be converted into a disparity map by combining the device parameters of the ranging device, so as to obtain the real disparity map corresponding to the training video frame.
[0165] During training of the disparity map generation model, the training video frame is first input into the disparity map generation model to obtain the disparity map output by the disparity map generation model as a predicted disparity map corresponding to the training video frame. Then, the predicted loss of the disparity map generation model is determined according to the predicted disparity map corresponding to the training video frame and the real disparity map corresponding to the training video frame. Finally, the parameters of the disparity map generation model are updated according to the predicted loss of the disparity map generation model. Different training video frames are used to iteratively train the disparity map generation model in the above manner until a training end condition is met (such as reaching a preset training number, or the performance of the disparity map generation model meeting the requirements).
[0166] The predicted loss of the disparity map generation model can be calculated based on the following formula:
[0167]
[0168] In the above formula, y in(i,j) represents the real disparity map corresponding to the input training video frame, y out (i,j) is the disparity map output by the disparity map generation model, i.e., the predicted disparity map corresponding to the training video frame.
[0169] Step S403: Determine the depth map corresponding to the video frame I based on the disparity map corresponding to the video frame I.
[0170] After obtaining the disparity map corresponding to the video frame I, the disparity map corresponding to the video frame I can be converted into a depth map in combination with the device parameters of the ranging device.
[0171] In another embodiment of the present application, the specific implementation process of "Step S204: locating the target waterlogging area in the target video frame according to the preliminary waterlogging area in each video frame included in the target video frame set and the depth map corresponding to each video frame in the target video frame set" in the above-mentioned embodiment is introduced.
[0172] Please refer to Figure 6 , which shows a flowchart of locating the target waterlogging area in the target video frame according to the preliminary waterlogging area in each video frame included in the target video frame set and the depth map corresponding to each video frame in the target video frame set, which can include:
[0173] Step S601: Correct the preliminary waterlogging area in each video frame included in the target video frame set according to the reference depth map and the depth map corresponding to each video frame in the target video frame set, to obtain the corrected waterlogging area in each video frame included in the target video frame set.
[0174] For each video frame included in the target video frame set, the preliminary waterlogging area in the video frame is corrected according to the reference depth map and the depth map corresponding to the video frame, to obtain the corrected waterlogging area in the video frame. The purpose of correcting the preliminary waterlogging area in each video frame included in the target video frame set is to exclude the pixel points that do not belong to the waterlogging area but are included in the waterlogging area.
[0175] The reference depth map is determined based on a reference video frame set, and the reference video frame set includes video frames without waterlogging areas obtained from the video under the specified video point.
[0176] Optionally, the process of obtaining the video frame without the waterlogging area from the video at the specified video point can include: obtaining a video segment in a sunny scene from the video at the specified video point; performing waterlogging area identification on each video frame contained in the video segment in the sunny scene respectively by using a semantic segmentation-based method to obtain identification results respectively corresponding to each video frame contained in the video segment in the sunny scene (the identification results can indicate whether there is a waterlogging area in the corresponding video frame); and determining the video frame without the waterlogging area from each video frame contained in the video segment in the sunny scene according to the identification results respectively corresponding to each video frame contained in the video segment in the sunny scene.
[0177] After obtaining the reference video frame set, the reference depth map can be determined based on the reference video frame set. Specifically, first, the depth map respectively corresponding to each video frame contained in the reference video frame set is obtained (the depth map respectively corresponding to each video frame contained in the reference video frame set can be obtained by using the method for obtaining the depth map corresponding to the video frame I provided in the above embodiment), and then the depth maps respectively corresponding to each video frame contained in the reference video frame set are averaged (the pixel values of the pixel points at the same position in the depth maps respectively corresponding to each video frame contained in the reference video frame set are averaged), to obtain an average depth map, which is taken as the reference depth map.
[0178] When the pixel values of the pixel points at the same position in the depth maps respectively corresponding to each video frame contained in the reference video frame set are averaged, in one possible implementation, the pixel values of all the pixel points at the same position in the depth maps respectively corresponding to each video frame contained in the reference video frame set can be directly averaged. For example, if the reference video frame set contains 100 video frames, the pixel values of 100 pixel points at the same position in the depth maps respectively corresponding to the 100 video frames are averaged. In order to reduce the interference of some interfering objects in the reference video frame set, in another possible implementation, the pixel values of all the pixel points at the same position in the depth maps respectively corresponding to each video frame contained in the reference video frame set can be sorted by size first, and then the first M (for example, 10) pixel values and the last M pixel values are removed, and the remaining pixel values are averaged. For example, if the reference video frame set contains 100 video frames, the pixel values of 100 pixel points at the same position in the depth maps respectively corresponding to the 100 video frames are sorted, the first 10 pixel values and the last 10 pixel values are removed, and the remaining 80 pixel values are averaged.
[0179] It should be noted that the reference video frame set can include one video frame without the water accumulation area or multiple video frames without the water accumulation area. If the reference video frame set includes one video frame without the water accumulation area, the depth map corresponding to the one video frame is obtained, and the depth map corresponding to the one video frame is directly used as the reference depth map. If the reference video frame set includes multiple video frames without the water accumulation area, the depth maps corresponding to the multiple video frames without the water accumulation area are averaged, and the obtained average depth map is used as the reference depth map. In order to determine a more accurate water accumulation area in the subsequent process, the reference video frame set preferably includes multiple video frames without the water accumulation area, and the average depth map obtained by averaging the depth maps corresponding to the multiple video frames without the water accumulation area is used as the reference depth map.
[0180] Since the way of correcting the preliminary water accumulation area in each video frame included in the target video frame set is the same, the embodiment takes one video frame I in the target video frame set as an example to introduce the specific implementation process of correcting the preliminary water accumulation area in the video frame I according to the reference depth map and the depth map corresponding to the video frame I.
[0181] The specific implementation process of correcting the preliminary water accumulation area in the video frame I according to the reference depth map and the depth map corresponding to the video frame I includes: performing the following steps for each pixel point included in the preliminary water accumulation area in the video frame I.
[0182] In step c1, the depth value corresponding to the pixel point is obtained from the depth map corresponding to the video frame I, and the reference depth value corresponding to the pixel point is obtained from the reference depth map.
[0183] Since the depth map corresponding to the video frame I includes the depth values respectively corresponding to the pixel points in the video frame I, the depth value corresponding to the pixel point can be obtained from the depth map corresponding to the video frame I. For example, the pixel point is the pixel point in the 10th row and the 8th column of the video frame I, and the pixel value of the pixel point in the 10th row and the 8th column of the depth map corresponding to the video frame I is the depth value corresponding to the pixel point. Therefore, when the depth value corresponding to the pixel point is obtained from the depth map corresponding to the video frame I, the pixel value of the pixel point in the 10th row and the 8th column of the depth map corresponding to the video frame I can be obtained. Similarly, the reference depth value corresponding to the pixel point can be obtained from the reference depth map.
[0184] In step c2, the difference between the depth value corresponding to the pixel point and the reference depth value corresponding to the pixel point is calculated as the depth difference value corresponding to the pixel point.
[0185] In step c3, it is determined whether the depth difference value corresponding to the pixel point is less than a preset water depth threshold. If yes, it is determined that the pixel point does not belong to the water accumulation area, and the pixel point is excluded from the preliminary water accumulation area.
[0186] Through the above process, the pixel points determined not to belong to the accumulated water region can be excluded from the preliminary accumulated water region, so as to obtain the corrected accumulated water region.
[0187] Step S602: determining the intersection region and the union region of the corrected accumulated water region in each video frame contained in the target video frame set.
[0188] For example, the target video frame set contains three 10*10 video frames, and the corrected accumulated water regions in the three video frames are as shown in FIG. 7(a) and FIG. 7(b). Figures 7(a) to 7(c) The intersection region of the corrected accumulated water regions in the three video frames is as shown in FIG. 8(a), and the union region of the corrected accumulated water regions in the three video frames is as shown in FIG. 8(b).
[0189] Step S603: locating the target accumulated water region in the target video frame according to the intersection region, the union region and the depth map corresponding to the target video frame.
[0190] Specifically, the process of locating the target accumulated water region in the target video frame according to the intersection region, the union region and the depth map corresponding to the target video frame can include:
[0191] Step S6031: obtaining the depth values corresponding to each pixel point respectively contained in the first accumulated water region and the second accumulated water region in the target video frame from the depth map corresponding to the target video frame.
[0192] The first accumulated water region is the region in the target video frame corresponding to the intersection region, and the second accumulated water region is the region in the target video frame corresponding to the union region.
[0193] For example, the intersection region shown in FIG. 8(a) contains pixel points at positions P 35 ~P 36 , P 44 ~P 47 , P 54 ~P 57 , P 64 ~P 66 , P 75 ~P 76 , wherein P 35 represents the 3rd row and the 5th column, and the first accumulated water region in the target video frame is the region in the target video frame containing P 35 ~P 36 , P 44 ~P 47 , P 54 ~P 57 , P 64 ~P 66 , P 75 ~P 76The region composed of the pixel points at the positions shown in Fig. 8(b) contains the pixel points at positions P 35 ~P 36 , P 44 ~P 47 , P 53 ~P 57 , P 63 ~P 67 , P 74 ~P 76 The second water area in the target video frame is a region composed of the pixel points at positions P 35 ~P 36 , P 44 ~P 47 , P 53 ~P 57 , P 63 ~P 67 , P 74 ~P 76 in the target video frame.
[0194] Step S6032: The average depth value of the first water area is obtained by averaging the depth values corresponding to the respective pixel points included in the first water area.
[0195] The sum of the depth values corresponding to the respective pixel points included in the first water area is obtained, and the obtained sum is divided by the number of the pixel points included in the first water area to obtain the average depth value of the first water area.
[0196] Step S6033: The depth difference values corresponding to the respective pixel points included in the second water area are obtained by calculating the difference between the depth values corresponding to the respective pixel points included in the second water area and the average depth value.
[0197] Step S6034: The pixel points included in the second water area, for which the corresponding depth difference values are greater than the preset depth change threshold, are determined as non-target pixel points.
[0198] For each pixel point included in the second water area, it is determined whether the depth difference value corresponding to the pixel point is greater than the preset depth change threshold. If yes, the pixel point is determined as a non-target pixel point. The non-target pixel point is a pixel point not belonging to the water area.
[0199] Step S6035: The region composed of the pixel points other than the non-target pixel points in the second water area is determined as the target water area in the target video frame.
[0200] For example, the second water area in the target video frame is a region composed of the pixel points at positions P 35 ~P 36 , P 44 ~P 47, P 53 ~ P 57 , P 63 ~ P 67 , P 74 ~ P 76 The pixel points in the position P form a region, and it is assumed that the pixel points in the position P 53 , P 74 The pixel points in the position P are non-target pixel points, and finally the pixel points in the position P 35 ~ P 36 , P 44 ~ P 47 , P 54 ~ P 57 , P 63 ~ P 67 , P 75 ~ P 76 The pixel points in the position P form a region, and it is assumed that the pixel points in the position P
[0201] The water accumulation region in the target video frame can be accurately positioned through the above process. After the water accumulation region in the target video frame is accurately positioned, the depth information of the water accumulation region can be obtained from the depth map corresponding to the target video frame, that is, the pixel value of the region corresponding to the positioned water accumulation region in the depth map corresponding to the target video frame is determined as the depth information of the positioned water accumulation region.
[0202] The water accumulation condition detection method provided by the present application first performs preliminary detection on the water accumulation region to determine the approximate range of the water accumulation region, and then further determines the accurate range of the water accumulation region in combination with the depth information, that is, the accurate water accumulation region is positioned in combination with the depth information, and on this basis, the depth information of the water accumulation region can be determined. Since the water accumulation condition detection method provided by the present application determines the water accumulation region and the water depth by analyzing the video (such as the video collected based on a monocular camera), without using auxiliary equipment or manual on-site measurement, the labor cost and the equipment cost caused by additional equipment are saved, and the restrictive problem caused by using auxiliary equipment to measure the depth is solved.
[0203] The embodiment of the present application also provides a water accumulation condition detection device, and the water accumulation condition detection device provided by the embodiment of the present application will be described below. The water accumulation condition detection device described below can be correspondingly referred to the water accumulation condition detection method described above.
[0204] Please refer to Figure 9 , which shows the structure schematic diagram of the water accumulation condition detection device provided by the embodiment of the present application. The water accumulation condition detection device can include a target video segment acquisition module 901, a water accumulation region identification module 902, a depth map acquisition module 903, a target water accumulation region determination module 904 and a water depth information determination module 905.
[0205] The target video segment acquisition module 901 is configured to acquire a target video segment at a specified video point.
[0206] The waterlogged area identification module 902 is configured to identify a waterlogged area in each video frame included in the target video segment as a preliminary waterlogged area.
[0207] The depth map acquisition module 903 is configured to acquire a depth map corresponding to each video frame included in a target video frame set, wherein the video frames included in the target video frame set are the video frames in which the waterlogged areas are identified.
[0208] The target waterlogged area determination module 904 is configured to determine a target waterlogged area in a target video frame from the preliminary waterlogged areas in each video frame included in the target video frame set and the depth map corresponding to each video frame in the target video frame set, wherein the target video frame is one video frame in the target video frame set.
[0209] The waterlogged depth information determination module 905 is configured to determine a pixel value of a region corresponding to the target waterlogged area in the depth map corresponding to the target video frame as depth information of the target waterlogged area.
[0210] Optionally, the target waterlogged area determination module 904 can include a waterlogged area correction sub-module, an intersection region and union region determination sub-module, and a target waterlogged area positioning sub-module.
[0211] The waterlogged area correction sub-module is configured to, for each video frame included in the target video frame set: correct the preliminary waterlogged area in the video frame based on a reference depth map and a depth map corresponding to the video frame to obtain a corrected waterlogged area in the video frame.
[0212] The reference depth map is determined based on a reference video frame set, and the reference video frame set includes video frames without waterlogged areas acquired from the video at the specified video point.
[0213] The intersection region and union region determination sub-module is configured to determine an intersection region and a union region of the corrected waterlogged areas in each video frame included in the target video frame set.
[0214] The target waterlogged area positioning sub-module is configured to determine the target waterlogged area in the target video frame from the intersection region, the union region, and the depth map corresponding to the target video frame.
[0215] Optionally, the waterlogged area correction sub-module, when correcting the preliminary waterlogged area in the video frame based on the reference depth map and the depth map corresponding to the video frame, is specifically configured to:
[0216] For each pixel point contained in the preliminary waterlogging area in the video frame:
[0217] A depth value corresponding to the pixel point is obtained from a depth map corresponding to the video frame, and a reference depth value corresponding to the pixel point is obtained from the reference depth map;
[0218] A difference value between the depth value corresponding to the pixel point and the reference depth value corresponding to the pixel point is calculated as a depth difference value corresponding to the pixel point;
[0219] If the depth difference value corresponding to the pixel point is less than a preset waterlogging depth threshold, it is determined that the pixel point does not belong to a waterlogging area, and the pixel point is excluded from the preliminary waterlogging area.
[0220] The waterlogging condition detection device provided by the embodiment of the application can further include a reference depth map determination module. The reference depth map determination module is configured to determine the reference depth map based on the reference video frame set.
[0221] Optionally, when determining the reference depth map based on the reference video frame set, the reference depth map determination module is specifically configured to:
[0222] Obtain depth maps corresponding to respective video frames contained in the reference video frame set;
[0223] Calculate average pixel values of pixel points at the same positions in the depth maps corresponding to the respective video frames contained in the reference video frame set, to obtain an average depth map, and the obtained average depth map is taken as the reference depth map.
[0224] Optionally, when locating the target waterlogging area in the target video frame based on the intersection area, the union area, and a depth map corresponding to the target video frame, the target waterlogging area positioning submodule is specifically configured to:
[0225] Obtain depth values corresponding to respective pixel points contained in the first waterlogging area and the second waterlogging area in the target video frame from a depth map corresponding to the target video frame, wherein the first waterlogging area and the second waterlogging area are areas corresponding to the intersection area in the target video frame and areas corresponding to the union area in the target video frame, respectively;
[0226] Calculate an average depth value by averaging the depth values corresponding to the respective pixel points contained in the first waterlogging area;
[0227] Calculate difference values between the depth values corresponding to the respective pixel points contained in the second waterlogging area and the average depth value, to obtain depth difference values corresponding to the respective pixel points contained in the second waterlogging area;
[0228] The pixel points in the second waterlogging area are determined as non-target pixel points if the corresponding depth values of the pixel points are greater than a preset depth change threshold value.
[0229] The region composed of the pixel points other than the non-target pixel points in the second waterlogging area is determined as a target waterlogging area in the target video frame.
[0230] Optionally, the waterlogging area identification module 902, when identifying the waterlogging areas in the video frames included in the target video segment, is specifically configured to:
[0231] For each video frame included in the target video segment:
[0232] predict the categories to which the pixel points included in the video frame respectively belong, wherein the category to which a pixel point belongs is one of waterlogging and background;
[0233] determine the region composed of the pixel points in the video frame whose category is waterlogging as the waterlogging area identified from the video frame.
[0234] Optionally, the target video segment is captured based on a monocular camera; and the depth map acquisition module 903 includes a disparity change map generation sub-module, a disparity map generation sub-module, and a depth map determination sub-module.
[0235] The disparity change map generation sub-module is configured to, for each video frame in the target video frame set: generate a disparity change map corresponding to the video frame based on the video frame. The disparity change map corresponding to the video frame is a right image generated for a left image of the video frame.
[0236] The disparity map generation sub-module is configured to generate a disparity map corresponding to the video frame based on the video frame and the disparity change map corresponding to the video frame.
[0237] The depth map determination sub-module is configured to determine a depth map corresponding to the video frame based on the disparity map corresponding to the video frame.
[0238] Optionally, the disparity change map generation sub-module, when generating the disparity change map corresponding to the video frame based on the video frame, is specifically configured to:
[0239] extract a plurality of levels of multi-channel feature maps from the video frame, and process the plurality of levels of multi-channel feature maps into a plurality of target feature maps with the same size as the video frame, wherein each target feature map includes feature information of the same channel at a plurality of levels;
[0240] predicting a disparity probability map under a set of disparity offset values based on the plurality of target feature maps, and respectively shifting positions of pixel values of the video frame based on the plurality of disparity offset values to obtain disparity offset maps under the plurality of disparity offset values;
[0241] generating a disparity change map corresponding to the video frame based on the disparity probability maps under the plurality of disparity offset values and the disparity offset maps under the plurality of disparity offset values.
[0242] Optionally, when the disparity change map generation sub-module generates the disparity change map corresponding to the video frame based on the plurality of disparity probability maps and the disparity offset maps respectively corresponding to the plurality of disparity probability maps, the disparity change map generation sub-module is specifically configured to:
[0243] multiply the disparity probability map and the disparity offset map under the same disparity offset value to obtain a plurality of multiplication results;
[0244] fuse the plurality of multiplication results, and take a fusion result as the disparity change map corresponding to the video frame.
[0245] Optionally, the disparity change map generation sub-module can adopt a disparity change map generation model trained in advance. The disparity change map generation model processes the video frame to generate the disparity change map corresponding to the video frame.
[0246] The disparity change map generation model is trained by taking a training video frame as a training sample and taking a real disparity change map corresponding to the training video frame as a sample label. The training video frame is a left image collected based on a binocular camera, and the real disparity change map corresponding to the training video frame is a right image collected based on the binocular camera.
[0247] Optionally, the disparity map generation sub-module can adopt a disparity map generation model trained in advance. The disparity map generation model processes the video frame and the disparity change map corresponding to the video frame to generate the disparity map corresponding to the video frame.
[0248] The disparity map generation model is trained by taking a training video frame and a training disparity change map corresponding to the training video frame as training samples and taking a real disparity map corresponding to the training video frame as a sample label.
[0249] The embodiment of the present application provides the waterlogging condition detection device, can automatically and accurately locate the waterlogging area in the video frame contained by the target video segment, and further determines the depth information of the located waterlogging area. When the waterlogging condition detection device provided by the embodiment of the present application realizes waterlogging condition detection, a measurement personnel does not need to hold a measuring device to measure the waterlogging depth of the waterlogging area, compared with the manual measurement mode, the measurement cost is greatly reduced, and the measurement efficiency is improved. In addition, when the waterlogging condition detection device provided by the embodiment of the present application realizes waterlogging condition detection, a water gauge is not needed, and a reference object is not needed, and has wide applicability.
[0250] The embodiment of the present application also provides a waterlogging condition detection device, please refer to Figure 10 , which shows the structural schematic diagram of the waterlogging condition detection device, and the waterlogging condition detection device can include: a processor 1001, a communication interface 1002, a memory 1003 and a communication bus 1004.
[0251] In the embodiment of the present application, the number of the processor 1001, the communication interface 1002, the memory 1003 and the communication bus 10010 is at least one, and the processor 1001, the communication interface 1002 and the memory 1003 complete mutual communication through the communication bus 1004.
[0252] The processor 1001 can be a central processing unit CPU, or a specific integrated circuit ASIC (Application Specific Integrated Circuit), or one or more integrated circuits configured to implement the embodiment of the present application, etc.
[0253] The memory 1003 can include a high-speed RAM memory, and can also include a non-volatile memory, etc., for example, at least one disk memory.
[0254] The memory stores a program, and the processor can call the program stored in the memory, and the program is used for:
[0255] Obtaining a target video segment under a specified video point;
[0256] Identifying the waterlogging area in each video frame contained by the target video segment as a preliminary waterlogging area;
[0257] Obtaining the depth map corresponding to each video frame in the target video frame set, wherein the video frame contained by the target video frame set is the video frame in which the waterlogging area is identified;
[0258] According to the preliminary waterlogging area in each video frame included in the target video frame set and the depth map corresponding to each video frame in the target video frame set, a target waterlogging area in a target video frame is located, wherein the target video frame is a video frame in the target video frame set.
[0259] A pixel value of a region corresponding to the target waterlogging area in the depth map corresponding to the target video frame is determined as the depth information of the target waterlogging area.
[0260] Optionally, the refinement function and the extension function of the program can refer to the description above.
[0261] The embodiment of the application further provides a readable storage medium which can store a program suitable for processor execution, and the program is used for:
[0262] obtaining a target video segment under a specified video point;
[0263] identifying a waterlogging area in each video frame included in the target video segment as a preliminary waterlogging area;
[0264] obtaining a depth map corresponding to each video frame in a target video frame set, wherein the video frame included in the target video frame set is a video frame in which the waterlogging area is identified;
[0265] According to the preliminary waterlogging area in each video frame included in the target video frame set and the depth map corresponding to each video frame in the target video frame set, a target waterlogging area in a target video frame is located, wherein the target video frame is a video frame in the target video frame set.
[0266] A pixel value of a region corresponding to the target waterlogging area in the depth map corresponding to the target video frame is determined as the depth information of the target waterlogging area.
[0267] Optionally, the refinement function and the extension function of the program can refer to the description above.
[0268] Finally, it should be noted that the terms "first", "second", and the like, herein do not denote any order, quantity, combination, or importance, but rather are used to distinguish one element from another, and are not intended to denote the presence of any such actual relationship or order. Moreover, the terms "comprises", "comprising", or any other variations thereof, are intended to cover a non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements does not include only those elements but can also include other elements not expressly listed or inherent to such process, method, article, or apparatus. An element proceeded by "comprises... a" does not, without more constraints, exclude the presence of additional identical elements in the process, method, article, or apparatus that comprises the element.
[0269] The various embodiments in the specification are described with progression in this order of description. Embodiments having the same or similar descriptions are referenced by the same reference numerals, and an overlapping description is not repeated.
[0270] The above description of disclosed embodiments provides enabling disclosure sufficient for one of ordinary skill in the art to practice the invention. Various modifications to these embodiments will be readily apparent to those skilled in the art, and generic principles defined herein can be applied to other embodiments without departing from the spirit or scope of the invention. Thus, the present invention is not intended to be limited to the embodiments shown herein but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A water-accumulation condition detection method characterized by comprising: The method comprises the following steps: acquiring a target video segment under a specified video point; identifying a water area in each video frame included in the target video segment as a preliminary water area; acquiring a depth map corresponding to each video frame included in a target video frame set, wherein the video frames included in the target video frame set are video frames in which the water area is identified; locating a target water area in a target video frame from the preliminary water area in each video frame included in the target video frame set and the depth map corresponding to each video frame in the target video frame set, wherein the target video frame is a video frame in the target video frame set; determining a pixel value of a region corresponding to the target water area in the depth map corresponding to the target video frame as depth information of the target water area; the step of locating the target water area in the target video frame from the preliminary water area in each video frame included in the target video frame set and the depth map corresponding to each video frame in the target video frame set comprises the following steps: for each video frame included in the target video frame set, correcting the preliminary water area in the video frame according to a reference depth map and a depth map corresponding to the video frame to obtain a corrected water area in the video frame, wherein the reference depth map is determined based on a reference video frame set, and the reference video frame set includes video frames without water area acquired from the video under the specified video point; determining an intersection region and a union region of the corrected water areas in each video frame included in the target video frame set; locating the target water area in the target video frame from the intersection region, the union region and the depth map corresponding to the target video frame; the step of locating the target water area in the target video frame from the intersection region, the union region and the depth map corresponding to the target video frame comprises the following steps: acquiring a depth value corresponding to each pixel point included in a first water area and a second water area in the target video frame from the depth map corresponding to the target video frame, wherein the first water area and the second water area are regions corresponding to the intersection region and the union region in the target video frame, respectively; averaging the depth values corresponding to each pixel point included in the first water area to obtain an average depth value; calculating a difference between the depth values corresponding to each pixel point included in the second water area and the average depth value to obtain a depth difference value corresponding to each pixel point included in the second water area; determining a pixel point as a non-target pixel point if the corresponding depth difference value of the pixel point is greater than a preset depth change threshold value; determining a region composed of other pixel points in the second water area except the non-target pixel points as the target water area in the target video frame.
2. The water-accumulation condition detection method according to claim 1, characterized by, the step of correcting the preliminary water area in the video frame according to the reference depth map and the depth map corresponding to the video frame comprises the following steps: for each pixel point included in the preliminary water area in the video frame, obtaining a depth value corresponding to the pixel point from a depth map corresponding to the video frame, and obtaining a reference depth value corresponding to the pixel point from the reference depth map; calculating a difference between the depth value corresponding to the pixel point and the reference depth value corresponding to the pixel point as a depth difference value corresponding to the pixel point; if the depth difference value corresponding to the pixel point is less than a preset water depth threshold, determining that the pixel point does not belong to the water area, and excluding the pixel point from the preliminary water area.
3. The water-accumulation condition detection method according to claim 1, characterized by, The process of determining the reference depth map based on the reference video frame set includes: obtaining the depth map corresponding to each video frame included in the reference video frame set; averaging the pixel values of the pixel points at the same position in the depth maps corresponding to each video frame included in the reference video frame set to obtain an average depth map, and taking the obtained average depth map as the reference depth map.
4. The water-accumulation condition detection method according to any one of claims 1 to 3, characterized by, The identification of the water area in each video frame included in the target video segment includes: for each video frame included in the target video segment: predicting the category to which each pixel point included in the video frame belongs, wherein the category to which a pixel point belongs is one of water and background; determining the region composed of the pixel points in the video frame whose category is water as the water area identified from the video frame.
5. The water-accumulation condition detection method according to any one of claims 1 to 3, characterized by, The target video segment is captured based on a monocular camera; The process of obtaining the depth map corresponding to each video frame included in the target video frame set includes: for each video frame in the target video frame set: generating a disparity change map corresponding to the video frame based on the video frame, wherein the disparity change map corresponding to the video frame is a right image generated based on the video frame as a left image; generating a disparity map corresponding to the video frame based on the video frame and the disparity change map corresponding to the video frame; determining the depth map corresponding to the video frame based on the disparity map corresponding to the video frame.
6. The water-accumulation condition detection method according to claim 5, wherein The process of generating the disparity change map corresponding to the video frame based on the video frame includes: extracting a plurality of hierarchical multi-channel feature maps from the video frame, and processing the plurality of hierarchical multi-channel feature maps into a plurality of target feature maps with the same size as the video frame, wherein each target feature map contains feature information of the same channel at multiple levels; based on the plurality of target feature maps, predicting disparity probability maps under a plurality of preset disparity offset values, and offsetting the positions of each pixel value of the video frame based on the plurality of disparity offset values to obtain disparity offset maps under the plurality of disparity offset values; generating the disparity change map corresponding to the video frame based on the disparity probability maps under the plurality of disparity offset values and the disparity offset maps under the plurality of disparity offset values.
7. The water-accumulation condition detection method according to claim 6, wherein The process of generating the disparity change map corresponding to the video frame based on the disparity probability maps under the plurality of disparity offset values and the disparity offset maps under the plurality of disparity offset values includes: multiplying the disparity probability map and the disparity offset map under the same disparity offset value to obtain a plurality of multiplication results; fusing the plurality of multiplication results, and taking the fusion result as the disparity change map corresponding to the video frame.
8. The water-accumulation condition detection method according to claim 5, wherein The process of generating the disparity change map corresponding to the video frame based on the video frame includes: The video frame is processed based on a disparity change map generation model trained in advance to generate a disparity change map corresponding to the video frame. The disparity change map generation model is trained based on training video frames and true disparity change maps corresponding to the training video frames, the training video frames are left images captured by a binocular camera, and the true disparity change maps are right images captured by the binocular camera.
9. The water-accumulation condition detection method according to claim 5, wherein The video frame and the disparity change map corresponding to the video frame are processed based on a disparity map generation model trained in advance to generate a disparity map corresponding to the video frame. The disparity map generation model is trained based on training video frames and true disparity maps corresponding to the training video frames. It comprises:
10. An accumulated water condition detection device characterized by comprising: a target video segment acquisition module, a water area identification module, a depth map acquisition module, a target water area determination module, and a water depth information determination module; The target video segment acquisition module is configured to acquire a target video segment at a specified video point. The water area identification module is configured to identify water areas in each video frame included in the target video segment as preliminary water areas. The depth map acquisition module is configured to acquire a depth map corresponding to each video frame included in a target video frame set, wherein the video frames included in the target video frame set are video frames in which water areas are identified. The target water area determination module is configured to locate a target water area in a target video frame based on preliminary water areas in each video frame included in the target video frame set and depth maps corresponding to the video frames in the target video frame set, wherein the target video frame is one of the video frames in the target video frame set. The water depth information determination module is configured to determine a pixel value of a region corresponding to the target water area in the depth map corresponding to the target video frame as depth information of the target water area. When locating the target water area in the target video frame based on the preliminary water areas in each video frame included in the target video frame set and the depth maps corresponding to the video frames in the target video frame set, the target water area determination module is specifically configured to: For each video frame included in the target video frame set, correct the preliminary water area in the video frame based on a reference depth map and a depth map corresponding to the video frame to obtain a corrected water area in the video frame, wherein the reference depth map is determined based on a reference video frame set, and the reference video frame set includes video frames without water areas acquired from the video at the specified video point. Determine an intersection region and a union region of the corrected water areas in the video frames included in the target video frame set. Obtaining, from a depth map corresponding to the target video frame, a depth value corresponding to each pixel point included in the first waterlogging area and the second waterlogging area in the target video frame respectively, wherein the first waterlogging area and the second waterlogging area are, in sequence, an area corresponding to the intersection area in the target video frame and an area corresponding to the union area in the target video frame; Obtaining an average depth value by averaging the depth values corresponding to each pixel point included in the first waterlogging area respectively; Obtaining a depth difference value corresponding to each pixel point included in the second waterlogging area by calculating a difference between the depth value corresponding to each pixel point included in the second waterlogging area and the average depth value; Determining, as a non-target pixel point, a pixel point corresponding to each pixel point included in the second waterlogging area and having a depth difference value greater than a preset depth change threshold value; Determining, as a target waterlogging area in the target video frame, an area formed by other pixel points in the second waterlogging area except the non-target pixel point.
11. An accumulated water condition detection apparatus characterized by comprising: Comprise: a memory and a processor; the memory is used to store a program; the processor is used to execute the program, and realize each step of the waterlogging condition detection method in any one of claims 1-9.
12. A computer readable storage medium having stored thereon a computer program, characterized in that, The computer program is executed by the processor to realize each step of the waterlogging condition detection method in any one of claims 1-9.
Citation Information
Patent Citations
Amniotic fluid depth measurement method and device, computer equipment and storage medium
CN117653206A