A method for repairing visually obscured areas of suspended impurities for underwater structure detection
By combining background displacement compensation and dynamic visual perception model with a hybrid restoration method, the problem of underwater suspended impurities occlusion is solved, high-quality image restoration of underwater structure detection is achieved, and detection accuracy and reliability are improved.
Patent Information
- Application Number
- CN202310936084.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-07-28
- Publication Date
- 2025-09-12
- Estimated Expiration
- 2043-07-28
AI Technical Summary
Existing technologies cannot effectively solve the problem of occlusion caused by multi-form underwater suspended impurities, which affects the accuracy and image quality of underwater structure detection.
Through background displacement compensation and dynamic visual perception model, combined with hybrid restoration methods, camera motion interference is eliminated, suspended impurity occlusion areas are detected and repaired, and comprehensive restoration is performed using background redundant information of adjacent frames and video restoration network.
It improves the quality of underwater images, enables accurate detection of multi-form suspended impurities and effective repair of blocked areas, and enhances the accuracy and reliability of underwater structure detection.
Smart Images

Figure CN117058018B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer vision technology, and in particular to a method for repairing a visually blocked area of suspended impurities for underwater structure detection, and in particular to a method for repairing a visually blocked area of suspended impurities for underwater structure apparent state detection. Background Art
[0002] Underwater structures are prone to various defects such as cracking, falling apart, and leakage during long-term operation, which can reduce their safety and reliability. Promptly troubleshooting safety hazards in underwater structures without draining them is crucial to ensuring the stable development of my country's socio-economic development. Currently, using underwater robots equipped with visible light cameras to collect and analyze surface images of underwater structures has become a common method for detecting their apparent condition. However, due to the numerous interference factors in underwater imaging, underwater image quality can be significantly reduced, adversely affecting subsequent underwater structure condition detection.
[0003] Currently, researchers are primarily concerned with image degradation caused by absorption by the water medium and scattering by suspended particles in the water. This degradation primarily manifests as a decrease in image clarity and contrast, but does not affect the representation of image content. This problem can be addressed to some extent by adding light sources and shortening the detection distance, and image quality can also be further improved through underwater image processing techniques. Numerous researchers have conducted extensive research on underwater image color correction and image dehazing, achieving significant results. However, existing research often overlooks the issue of image information loss caused by visual occlusion caused by large suspended particles in real aquatic environments. Large suspended particles obstruct a large area, altering underwater content. Individual small suspended particles obstruct a small area, but when densely distributed within an image, they can severely degrade underwater images. In particular, when the robot is operating in water, the rotation and automatic movement of its propellers accelerates the water flow, dispersing suspended particles such as algae and rotting leaves. This further exacerbates visual occlusion, severely impacting subsequent information processing and the accurate detection of the apparent condition of underwater structures. Therefore, studying the method of eliminating underwater suspended impurities obstruction is of great significance to improving the detection effect of the apparent state of underwater structures.
[0004] Currently, there is relatively little research on the removal of underwater suspended impurities, primarily focusing on the removal of marine snow in the field of ocean exploration. Marine snow is a type of suspended impurity, primarily composed of smaller plants and animals in the ocean. Its particles resemble floating snowflakes in the water, hence the name marine snow. Because marine snow particles are mostly small in size and have an overall shape that approximates an elliptical cone, they form white bright spots in the image. Therefore, many researchers treat it as speckle noise in underwater images and process it accordingly. However, current research is based on the detection of a single visual characteristic of marine snow, which cannot meet the needs of removing occlusions from multi-form underwater suspended impurities. Research is needed on more adaptable methods for underwater suspended impurity detection and occluded area repair. Summary of the Invention
[0005] In response to the shortcomings of the existing technology, the present invention provides a method for repairing the visual occlusion area of suspended impurities for underwater structure detection, which is used for the research on the detection and removal tasks of underwater suspended impurities and can solve the problem of insufficient accuracy in eliminating occlusions of multi-form underwater suspended impurities.
[0006] To achieve the above object, the present invention provides the following technical solutions:
[0007] The method for repairing the visual occlusion area of suspended impurities for underwater structure detection includes the following steps:
[0008] Step S1: The current frame I k and its adjacent frames I k-1 with I k+1 The three-frame segment composed of the current frame is used as input data, and the motion optical flow field between the current frame and the two previous and next frames is obtained respectively. According to the motion optical flow field information and the edge-guided region segmentation information, the background displacement value between adjacent frames is calculated, and the background displacement is used as compensation information to align the adjacent frames to eliminate the interference of camera motion on the detection of suspended impurity motion information. The current frame I k , aligned previous adjacent frames and the aligned adjacent frames
[0009] According to the optical flow field distribution information between adjacent frames, a background displacement compensation strategy is proposed to eliminate the background offset between frames caused by camera movement, that is, to eliminate the interference of camera motion on the detection of suspended impurity motion information.
[0010] Step S2: Using inter-frame differences to establish a dynamic visual perception model, extract the aligned previous adjacent frames obtained in step S1 Aligned with the next adjacent frame The motion and color information between the adjacent frames is used to complement each other to obtain the suspended impurity detection result area S k ;
[0011] Combining the imaging characteristics of suspended impurities, a dynamic visual perception model is established to accurately detect suspended impurities of different forms based on aligned adjacent frames.
[0012] It should be noted that "motion and color information" refers to the inherent information of aligned adjacent frames. "Motion information" uses the frame difference method to obtain the motion information of suspended impurities in the adjacent frame and the current frame. "Color information" is obtained by combining the Canny and UCM algorithms, taking into account the characteristic that the color of suspended impurities and the background of underwater structures are generally significantly different.
[0013] Step S3: Through the suspended impurities detection result area S k , that is, the suspended impurity occlusion area is obtained, which is used to repair the suspended impurity occlusion area; the aligned previous adjacent frame obtained in step S1 is used Aligned with the next adjacent frame The background redundant information between the adjacent frames (based on the information obtained from the adjacent frames that are not blocked by suspended impurities in the suspended impurity occlusion area of the current frame, the adjacent frames have areas that are not blocked by suspended impurities in the block area, that is, the background redundant information, which can be used to preliminarily repair the partial area blocked by suspended impurities in the current frame), build a hybrid repair model, determine the best matching area in the aligned adjacent frames, perform the first repair on the suspended impurity occlusion area of the current frame, and obtain the first repair result I k , and then combined with the joint space-time transformation network STTN of video restoration to perform a second restoration on the unrepairable area, and obtain the second restoration result Fusion of two restoration results and Get the final repaired image
[0014] Construct a hybrid restoration model, establish the optimal complementary information between frames, and restore the area occluded by suspended impurities.
[0015] Step S4: After the current input data processing is completed, three frame segments are updated, and the processing from step S1 to step S4 is repeated until all video frames are restored.
[0016] Furthermore, the step S1 specifically includes the following steps:
[0017] Step S11: Use Selflow optical flow method to obtain the current frame I k and the adjacent frames I k-1 with I k+1 The motion optical flow field is converted into a visual color map of the adjacent frames. and
[0018] Step S12: Using the Canny algorithm to visualize the color images of the adjacent frames obtained in step S11 and Perform edge detection, calculate the edge probability of each pixel, perform directional watershed transform and ultra-high-depth contour map transform (UCM) on the edge map, convert the open boundary probability area into several closed areas, and obtain the closed areas of the adjacent frames. and
[0019] Step S13: The closed regions of the adjacent frames obtained in step S12 are respectively and The mean of the image segmentation threshold is used to obtain the visual color map of the adjacent frames before and after and The region segmentation result and
[0020] Step S14: Visualize the optical flow color map of the adjacent frames obtained in step S11 and Convert to CIELAB color space, construct the color histogram in each area, and calculate the weighted contrast Ctr between the segmented area i and other areas i c , get the foreground area probability estimation map of the adjacent frames before and after and
[0021]
[0022] Among them, N is the number of regions obtained by segmentation according to the image segmentation threshold; i and j are both segmented regions, and both are and The area in, i = 1, 2, 3, ..., N; and Represents the color histogram of segmented regions i and j, j = 1, 2, 3, ..., N, i ≠ j, ||·||2 is the Euclidean distance, ||p i ,p j ||2 is the distance between the center points of segmented regions i and j, σ spa is the parameter of the spatial weighting scheme, and the empirical value here is 0.4;
[0023] Step S15: Use the maximum inter-class variance method to calculate the foreground region probability estimation map of the adjacent frames obtained in step S14 and The global image threshold is used to obtain the distinguishing mark map of the foreground and background areas of the adjacent frames. and
[0024] Step S16: Compensate the foreground area according to the background displacement vector surrounding the foreground area. The foreground area surrounding refers to the 8-connected domain window centered at the edge pixel point of the dilated area in the following S16a.
[0025] Furthermore, the specific steps of step S16 are:
[0026] Step S16a: transform the previous adjacent frame so that its background area is consistent with the background area of the current frame, and extract any foreground area Rf i , and perform expansion processing on it, and construct 8 connected domain windows with the edge pixels of the expansion area as the center in turn, calculate the mean of the optical flow vectors in all windows as the background displacement compensation value of the foreground area, and cyclically calculate the background displacement of all foreground areas as the compensation information of the foreground target area to obtain the background optical flow field after compensation of the previous adjacent frame Then follow the same steps to obtain the background optical flow field after compensation of the adjacent frames
[0027]
[0028] Among them, x′ represents the edge pixel point of the foreground target area after expansion, Indicates the average optical flow vector value of the 8-connected domain corresponding to the edge pixel x′ of the previous adjacent frame, It represents the average optical flow vector value of the 8-connected domain corresponding to the edge pixel point x′ in the adjacent frame, and mean(·) is the mean calculation symbol;
[0029] Step S16b: Using the background optical flow field after compensation of the previous adjacent frame obtained in step S16a Background optical flow field after compensation with the adjacent frame Perform alignment transformation on the front and back adjacent frames respectively, that is, adjust the image coordinates accordingly according to the optical flow vector to obtain the aligned front adjacent frames Aligned with the next adjacent frame
[0030]
[0031] Wherein, x represents the horizontal coordinate of the edge pixel point, and y represents the vertical coordinate of the edge pixel point.
[0032] Furthermore, the step S2 specifically includes the following steps:
[0033] Step 21: For the current frame I k and the previous adjacent frame aligned with the one obtained in step S1 and aligned adjacent frames Calculate the forward frame difference and the backward frame difference to obtain the difference map between the current frame and the previous adjacent frame The difference map between the current frame and the next adjacent frame And use it as motion information for dynamic perception:
[0034]
[0035] Where (x,y) represents and The pixel position in , T′ is the difference image threshold, which is used to filter the noise of the difference image in the three channels of RGB respectively. In this invention, the empirical value is 15;
[0036] Step 22: The difference image between the current frame and the previous adjacent frame obtained in step 21 The difference map between the current frame and the next adjacent frame Perform median filtering to reduce the impact of imaging noise, illumination changes and other factors on the difference image and obtain the forward motion feature map and backward motion feature map
[0037] Step S23: Use the maximum inter-class variance method to obtain the forward motion feature map Threshold and backward motion feature map Threshold
[0038]
[0039] Among them, Otsu(·) represents the global image threshold obtained by the maximum inter-class variance method, mean(·) represents the calculated mean, and the small-scale suspended impurity detection image is obtained by full-value threshold segmentation. and "Small-sized suspended impurities" are defined as areas smaller than 1 / 10 of the entire image. All other suspended impurities are considered large-sized. Large-sized suspended impurities are defined as areas no smaller than 1 / 10 of the entire image.
[0040] Step 24: Use the method of combining Canny and UCM to obtain the current frame I k The region segmentation result Realize the distinction between the background area and the large-scale suspended impurity area; the method of combining Canny and UCM is as follows: consistent with step S12, specifically, first use the Canny algorithm to detect the edge of the current frame to obtain the edge probability of each point, then perform directional watershed transformation and super-degree contour map transformation UCM on the edge map, convert the open boundary probability area into several closed areas, and obtain the area segmentation result of the current frame.
[0041] Step 25: Extract the region segmentation results obtained in step 24 in sequence For each pixel point in the image, we calculate the mean of the forward motion feature value and the backward motion feature value in each area to generate a motion contrast perception heat map. and Then, the refined landmark map of forward motion is obtained by threshold segmentation and the refined landmark map of backward motion The k in the is the k-th frame image or the k-th three-frame segment, and its specific value depends on the length of the video frame.
[0042] Step 26: Compute the refined landmark map for forward motion and the refined landmark map of backward motion Motion contrast perception heat map obtained based on global optical flow field and The intersection of the two results yields the refinement of the large-scale suspended impurity region. and Right now
[0043] Step 27: Small size suspended impurity detection map and Refinement results of large-scale suspended impurity detection area and Combine, that is, take the union, to get the final detection result of suspended impurities
[0044] Furthermore, the step S3 specifically includes the following steps:
[0045] Step 31: Construct a hybrid restoration label map based on the best matching area in the adjacent frame, represented by L = {0, 1, 2, 3}, where L(x, y) = 1 indicates that the pixel is not blocked by floating impurities and does not need to be restored; when L(x, y) = 2, it means that the best matching point of the pixel is the previous adjacent frame. When L(x,y)=3, it means that the best matching point of the pixel is the next adjacent frame. When L(x,y)=0, it means that no valid matching point is found for the pixel in the three-frame segment for repair, and it is necessary to estimate it by learning the surrounding information;
[0046] Step 32: Using the method of minimizing the label cost, solve the mixed repair label graph L, and calculate the minimized label cost C(L) according to the matching conditions in step 31:
[0047]
[0048] Among them, C d (L) is used to measure the difference between the aligned adjacent frames and the current frame; C s (L) is used to measure the smoothing cost between the best matching region in the aligned adjacent frames and the neighboring region in the current frame (the neighboring region refers to the region of the 4-connected neighbors described below); τ is the regularization parameter, τ = 50; Indicates the aligned adjacent frames, including the aligned previous adjacent frames Aligned with the next adjacent frame is the set of 4-connected neighbors of the pixel point (x,y); Is the exclusive OR operator; x and y represent Pixels within the region;
[0049] Step 33: By minimizing the label cost C(L), the hybrid repair label map L is obtained, and then the area in the current frame that is not blocked by suspended impurities (i.e., the mask S k =0) set all to 0.
[0050] Furthermore, the step S33 specifically includes the following steps:
[0051] Step 33A: Based on the hybrid restoration label map L, the current frame I k Perform the first restoration and use the multi-band fusion method to eliminate the uneven illumination traces of the restored image. The specific process is as follows;
[0052] Step 33A1: Use Gaussian pyramid to restore the current frame I k , aligned previous adjacent frames and the next adjacent frame Decomposed into multiple multi-band sub-images, respectively, with I k As the bottom image of the Gaussian pyramid Then pass it through a 5×5 low-pass Gaussian kernel g 5×5 From the underlying image Start convolution and downsampling layer by layer to obtain other pyramid layer images, and finally build an N ga +1 Gaussian pyramid image layer dw(a,b) means downsampling a by b times, i′ means the i′th layer Gaussian pyramid image, N ga Indicates the number of Gaussian pyramid image layers;
[0053] Step 33A2: In order to reduce redundant information in the Gaussian pyramid, it is necessary to further construct a Laplacian pyramid. The Laplacian image layer is obtained by subtracting two adjacent layers of the Gaussian pyramid from bottom to top:
[0054]
[0055] in, Represents the Laplacian pyramid image layer;
[0056] Select the part L={1,2,3} in the hybrid repair label map as the fusion mask and construct the combined pyramid:
[0057]
[0058] Step 33A3: Upsample the combined pyramid of each layer to the original image size and accumulate to obtain the first restoration result
[0059]
[0060] Step 33B: For the areas in the adjacent frames that cannot provide effective repair information, the joint space-time transform network STTN of video repair is used for further repair. The three frame segments are used as network input data, and the suspended impurity detection result area S k As an occlusion mask, perform a second repair on the current frame to obtain the second repair result
[0061] Step 33C: Target Part of the structural information of the repaired area is blurred, resulting in the loss of internal details. and Perform fusion to obtain the final repaired image
[0062]
[0063] Wherein, v′ is a set weight coefficient, which is 0.5 in the present invention.
[0064] Compared with the existing technology, the present invention provides a method for repairing the visual obstruction area caused by suspended impurities for underwater structure detection, which has the following beneficial effects:
[0065] (1) The suspended impurity occlusion elimination method of the present invention is based on the alignment of adjacent frames with background displacement compensation, which can effectively restore the background area of adjacent frames and avoid the interference of the relative motion between the foreground and background areas. It can solve the problem of insufficient accuracy in eliminating occlusions of multi-form underwater suspended impurities, effectively improve the quality of underwater images, and has high engineering value and application value.
[0066] (2) The underwater suspended impurity area detection based on dynamic visual perception proposed by the suspended impurity occlusion elimination method of the present invention uses the foreground area as a rough detection result of suspended impurities, and adopts inter-frame difference to construct a dynamic visual perception model to supplement and refine the suspended impurity area, thereby realizing comprehensive and accurate detection of suspended impurity targets of different sizes.
[0067] (3) The suspended impurity occlusion area repair guided by the hybrid repair model proposed in the suspended impurity occlusion elimination method of the present invention utilizes the background redundant information between the aligned adjacent frames to construct a hybrid repair model, determine the best matching area in the aligned adjacent frames, and perform a preliminary restoration on the suspended impurity occlusion area of the current frame. Then, the joint spatiotemporal transformation network of video repair is combined to supplement the area that cannot be repaired, which not only retains the real background information but also can better and completely repair the image. BRIEF DESCRIPTION OF THE DRAWINGS
[0068] Figure 1 This is a flow chart of the suspended impurity occlusion elimination method of the present invention;
[0069] Figure 2 A schematic diagram of the adjacent frame alignment process based on background displacement compensation provided by the present invention;
[0070] Figure 3 A schematic diagram of the underwater suspended impurity area detection process of the dynamic visual perception model provided by the present invention;
[0071] Figure 4 This is a flow chart of repairing suspended impurity-occluded areas guided by the hybrid repair model provided by the present invention. DETAILED DESCRIPTION
[0072] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.
[0073] Unless otherwise specifically stated, the relative arrangement, numerical expressions and numerical values of the parts and steps set forth in these embodiments do not limit the scope of the present invention. Meanwhile, it should be understood that, for ease of description, the sizes of the various parts shown in the accompanying drawings are not drawn according to actual proportional relationships. Technology, methods and equipment known to those of ordinary skill in the relevant art may not be discussed in detail, but in appropriate cases, the technology, methods and equipment should be considered as a part of the specification. In all examples shown and discussed here, any specific value should be interpreted as being merely exemplary, rather than as a limitation. Therefore, other examples of exemplary embodiments may also include different values. It should be noted that similar numbers and letters represent similar items in the following drawings, and therefore, once an item is defined in an accompanying drawing, it does not need to be further discussed in subsequent drawings.
[0074] In the description of this application, it should be understood that the terms "center", "longitudinal", "lateral", "up", "down", "left", "right", "vertical", "horizontal", "top", "bottom", "inside", "outside", etc., indicating the orientation or position relationship, are based on the orientation or position relationship shown in the accompanying drawings, and are only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore cannot be understood as a limitation on the protection content of the present invention.
[0075] The present invention discloses a method for repairing visually obscured areas caused by suspended impurities for underwater structure detection. Compared with traditional robot navigation and obstacle avoidance methods, the method improves the robot's ability to perceive obstacles, thereby further improving the safety of the cleaning robot's actual navigation. The path planning method has higher search efficiency, higher accuracy and sensitivity of the planned path, and is more adaptable to the cleaning needs of photovoltaic power station arrays.
[0076] like Figure 1 As shown, the method for eliminating suspended impurities obstruction in underwater structure status observation proposed by the present invention includes the following steps:
[0077] Step 1: Set the current frame I k and its adjacent frames I k-1 with I k+1 The three-frame segment is used as input data, such as Figure 2 As shown in the figure, the motion optical flow field between the current frame and the two previous and next frames is obtained respectively. According to the motion optical flow field information and the edge-guided region segmentation information, the background displacement value between adjacent frames is calculated, and the background displacement is used as compensation information to align the adjacent frames to eliminate the interference of camera motion on the detection of suspended impurity motion information. The current frame I is obtained. k , aligned previous adjacent frames and the aligned adjacent frames The specific steps are as follows:
[0078] 11) Use Selflow optical flow method to estimate the current frame I k and the adjacent frames I k-1 with I k+1 The motion optical flow field is further converted into a visual color map and
[0079] 12) Use the Canny algorithm to visualize the color images of the adjacent frames obtained in 11) and Perform edge detection and calculate the edge probability of each pixel. Let E(x,y,θ) represent the edge probability prediction value of the pixel at position (x,y) in the θ direction. Represents the maximum edge prediction value at the (x, y) position, performs directional watershed transformation and ultra-high-degree contour map transformation UCM on the edge map, and since the edge probability is mostly an open curve, the open boundary probability area is converted into several closed areas to obtain a closed area and The specific steps are as follows:
[0080] 12a) Select the minimum point in E(x,y) as the water injection point, and use the watershed transform to over-segment the edge map to obtain a series of catchment basin areas P0 and the arc boundaries K0 of the corresponding watersheds as the initial segmentation areas with rich details.
[0081] 12b) The arc-shaped boundary K0 is subdivided into approximate straight line segments K'0, and the average edge probability W(K'0) of each pixel on K'0 is calculated as the boundary strength.
[0082] 12c) A graph-based region merging algorithm is used to obtain multi-level region segmentation contours. Region P0 obtained by watershed transform is used as the vertex to construct a graph G = (P0, V, W(K'0)), where V represents the edge connecting each vertex. The average intensity value W(K'0) of the common boundary between two adjacent regions is set as the edge weight to measure the similarity between nodes. The most similar regions are iteratively merged. The region merging process here can be regarded as the growth of a "tree". The entire image is used as the root node of the tree, and region P0 is the leaf node of the tree. The height of each leaf node represents the segmentation threshold of each region. The boundary map is used to generate a set of hierarchical UCM edge probability maps to obtain closed regions. and
[0083] 13) The closed areas of the adjacent frames obtained in 12) are respectively and The mean of the image segmentation threshold is used to obtain the visual color map of the adjacent frames before and after and The region segmentation result and
[0084] 14) Visualize the optical flow color map of the adjacent frames obtained in 11) and Convert to CIELAB color space, construct the color histogram in each area, and calculate the weighted contrast between segmented area i and other areas Get the foreground region probability estimation map of the adjacent frames and
[0085]
[0086] Among them, N is the number of regions obtained by segmentation according to the image segmentation threshold; i and j are both segmented regions, and both are and The area in, i = 1, 2, 3, ..., N; and Represents the color histogram of segmented regions i and j, j = 1, 2, 3, ..., N, i ≠ j, ||·||2 is the Euclidean distance, ||p i ,p j ||2 is the distance between the center points of segmented regions i and j, σ spa is the parameter of the spatial weighting scheme, and the empirical value here is 0.4.
[0087] 15) Use the maximum inter-class difference method Otsu to calculate the foreground region probability estimation map of the adjacent frames obtained in 14) and The global image threshold is used to obtain the distinguishing mark map of the foreground and background areas of the adjacent frames. and
[0088] 16) Compensate the foreground area based on the background displacement vector surrounding the foreground area. The specific steps are as follows:
[0089] 16a) Transform the previous adjacent frame so that its background area is consistent with the background area of the current frame, and extract any foreground area Rf i , and perform expansion processing on it, and construct 8 connected domain windows with the edge pixels of the expansion area as the center in turn, calculate the mean of the optical flow vectors in all windows as the background displacement compensation value of the foreground area, and cyclically calculate the background displacement of all foreground areas as the compensation information of the foreground target area to obtain the background optical flow field after compensation of the previous adjacent frame Then follow the same steps to obtain the background optical flow field after compensation of the adjacent frames
[0090]
[0091] Among them, x′ represents the edge pixel point of the foreground target area after expansion, Indicates the average optical flow vector value of the 8-connected domain corresponding to the edge pixel x′ of the previous adjacent frame, It represents the average optical flow vector value of the 8-connected domain corresponding to the edge pixel point x′ in the adjacent frame, and mean(·) is the mean calculation symbol;
[0092] 16b) The background optical flow field after compensation of the previous adjacent frame obtained by 16a) Background optical flow field after compensation with the adjacent frame Perform alignment transformation on the front and back adjacent frames respectively, that is, adjust the image coordinates accordingly according to the optical flow vector to obtain the aligned front adjacent frames Aligned with the next adjacent frame
[0093]
[0094] Wherein, x represents the horizontal coordinate of the edge pixel point, and y represents the vertical coordinate of the edge pixel point.
[0095] Step 2: If Figure 3 As shown, the dynamic visual perception model is constructed using inter-frame differences, and the aligned previous adjacent frames obtained in step 1 are extracted. Aligned with the next adjacent frame The motion and color information between the adjacent frames is used to complement each other to obtain the suspended impurity detection result area S k The specific process is:
[0096] 21) For the current frame I k Aligned with the previous adjacent frame obtained in step 1 and aligned adjacent frames Calculate the forward frame difference and the backward frame difference to obtain the difference map between the current frame and the previous adjacent frame The difference map between the current frame and the next adjacent frame And use it as motion information for dynamic perception:
[0097]
[0098] Where (x,y) represents and The position of the pixel in , T′ is the difference image threshold, which is used to filter the noise of the difference image in the three channels of RGB respectively. In this invention, the empirical value is 15.
[0099] 22) The difference image between the current frame and the previous adjacent frame obtained in 21) The difference map between the current frame and the next adjacent frame Perform median filtering to reduce the impact of imaging noise, illumination changes and other factors on the difference image and obtain the forward motion feature map and backward motion feature map
[0100] 23) Use the maximum inter-class variance method to obtain the forward motion feature map and backward motion feature map Threshold and
[0101]
[0102] Among them, Otsu(·) represents the global image threshold obtained by the maximum inter-class variance method, mean(·) represents the calculated mean, and the small-scale suspended impurity detection image is obtained by full-value threshold segmentation. and If the suspended impurity area is less than 1 / 10 of the entire image, it is considered "small-sized suspended impurities".
[0103] 24) Use the same Canny and UCM method as 12) to obtain the current frame I k The region segmentation result Achieve differentiation between background area and large-sized suspended impurity area;
[0104] 25) Extract the region segmentation results obtained in 24) in sequence For each pixel point in the image, we calculate the mean of the forward motion feature value and the backward motion feature value in each area to generate a motion contrast perception heat map. and Then, the refined landmark map of forward motion is obtained by threshold segmentation and the refined landmark map of backward motion The k in the code refers to the kth frame image or the kth three-frame segment. The specific value depends on the length of the video frame. Figure 3 K in represents the current frame I k The region segmentation result There are K segmentation regions, and the value range is greater than or equal to 1.
[0105] 26) Calculate the refined landmark map for forward motion and the refined landmark map of backward motion Motion contrast perception heat map obtained based on global optical flow field and The intersection of the two methods obtains the refinement result of the large-scale suspended impurity detection area. and Right now
[0106] 27) Small size suspended impurities detection map and Refinement results of large-scale suspended impurity detection area and Combine, that is, take the union, to get the final detection result of suspended impurities
[0107] Step 3: If Figure 4 As shown, using the aligned previous adjacent frames Aligned with the next adjacent frame Based on the redundant background information between the two frames, a hybrid restoration model is constructed to determine the best matching area in the aligned adjacent frames. The suspended impurity occlusion area of the current frame is restored for the first time. Then, the joint spatiotemporal transform network (STTN) of video restoration is combined to perform supplementary restoration on the unrepairable areas, thus achieving comprehensive restoration of the suspended impurity occlusion area. The specific process is as follows:
[0108] 31) Construct a hybrid repair label map based on the best matching area in the adjacent frames, represented by L = {0, 1, 2, 3}, where L(x, y) = 1 means that the pixel is not blocked by floating impurities and does not need to be repaired; when L(x, y) = 2, it means that the best matching point of the pixel is the adjacent previous frame. When L(x,y)=3, it means that the best matching point of the pixel is the next frame. When L(x,y)=0, it means that no valid matching point is found for the pixel in the three-frame segment for repair, and it is necessary to estimate it by learning the surrounding information;
[0109] 32) Use the method of minimizing the label cost to solve the mixed repair label graph and calculate the minimized label cost C(L) according to the matching conditions in 31):
[0110]
[0111] Among them, C d (L) is used to measure the difference between the aligned adjacent frames and the current frame; C s (L) is used to measure the smoothing cost between the best matching region in the aligned adjacent frames and the adjacent region in the current frame; τ is the regularization parameter, τ = 50; Indicates the aligned adjacent frames, including the aligned previous adjacent frames Aligned with the next adjacent frame is the set of 4-connected neighbors of the pixel point (x,y), Is the exclusive OR operator; x and y represent Pixels within the region;
[0112] 33) By minimizing the label cost C(L), the mixed repair label map L is obtained, and then the area in the current frame that is not blocked by suspended impurities (i.e., mask S k =0) are all set to 0, the specific process is as follows:
[0113] 33A) Based on the hybrid restoration label map L, the current frame I k Perform preliminary restoration and use multi-band fusion method to eliminate uneven illumination traces in the restored image. The specific process is as follows;
[0114] 33a) Using Gaussian pyramid to restore the current frame I k , aligned previous adjacent frames and the next adjacent frame Decomposed into multiple multi-band sub-images, respectively, with I k As the bottom image of the Gaussian pyramid Then pass it through a 5×5 low-pass Gaussian kernel g 5×5 Starting from the bottom image, convolution and downsampling are performed layer by layer to obtain other pyramid layer images, and finally an N ga +1 Gaussian pyramid image layer dw(a,b) means downsampling a by b times, i′ means the i′th layer Gaussian pyramid image, N ga Indicates the number of Gaussian pyramid image layers;
[0115] 33b) In order to reduce the redundant information in the Gaussian pyramid, it is necessary to further construct a Laplacian pyramid. The Laplacian image layer is obtained by subtracting two adjacent layers of the Gaussian pyramid from bottom to top:
[0116]
[0117] in, Represents the Laplacian pyramid image layer;
[0118] Select the part L={1,2,3} in the hybrid repair label map as the fusion mask and construct the combined pyramid:
[0119]
[0120] 33c) Upsample each layer of the combined pyramid to the original image size and accumulate to obtain the first restoration result
[0121]
[0122] 33B) For the areas in the adjacent frames that cannot provide effective repair information, the joint space-time transformation network STTN of video repair is used for further repair. The three frame segments are used as network input data, and the suspended impurity detection result area S k As an occlusion mask, perform a second repair on the current frame to obtain the second repair result
[0123] 33C) for Part of the structural information of the repaired area is blurred, resulting in the loss of internal details. and Perform fusion to obtain the final repaired image
[0124]
[0125] Wherein, v′ is the set weight coefficient, which is 0.5 in the present invention;
[0126] Step 4: After the current input data is processed, update the three-frame segment until all video frames are repaired.
[0127] It should be noted that the joint space-time transformation network (STTN) for video restoration in this application is an existing technology and has not been improved. Figure 4 The symbols Q, K, V, t, h, w, r1, and r2 in the “STTN video frame restoration model” correspond to the usage meanings of the joint spatiotemporal transform network STTN for video restoration in the prior art.
[0128] It should be noted that, in this application, relational terms such as first and second, etc. are merely used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply the existence of any such actual relationship or order between these entities or operations. Moreover, the terms "comprises," "comprising," or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article, or device comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article, or device. In the absence of further restrictions, an element defined by the statement "comprising a ..." does not exclude the presence of other identical elements in the process, method, article, or device comprising the element.
[0129] While embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions, and variations may be made to these embodiments without departing from the principles and spirit of the invention, and that the scope of the invention is defined by the appended claims and their equivalents.
Claims
1. A method for repairing visually obscured areas caused by suspended impurities for underwater structure detection, characterized in that: The following steps are involved: Step S1: Set the current frame and its adjacent frames and The three-frame segment is used as input data to obtain the motion optical flow field between the current frame and the two frames before and after. According to the motion optical flow field information and the edge-guided region segmentation information, the background displacement value between adjacent frames is calculated, and the background displacement is used as compensation information to align adjacent frames to eliminate the interference of camera motion on the detection of suspended impurities motion information. The current frame is obtained. , aligned previous adjacent frames and the aligned adjacent frames ; Step S2: Using inter-frame differences to establish a dynamic visual perception model, extract the aligned previous adjacent frames obtained in step S1 Aligned with the next adjacent frame The motion and color information between the adjacent frames is used to complement each other to obtain the suspended impurity detection result area. ; Step S3: Through the suspended impurities detection result area , that is, the suspended impurity occlusion area is obtained, which is used to repair the suspended impurity occlusion area; the aligned previous adjacent frame obtained in step S1 is used Aligned with the next adjacent frame The background redundant information between the frames is used to build a hybrid repair model, determine the best matching area in the aligned adjacent frames, and perform the first repair on the suspended impurity occluded area of the current frame to obtain the first repair result. , and then combined with the joint space-time transformation network STTN of video restoration to perform a second restoration on the unrepairable area, and obtain the second restoration result , fusion of two restoration results and Get the final repaired image ; Step S4: After the current input data processing is completed, three frame segments are updated, and the processing from step S1 to step S4 is repeated until all video frames are restored.
2. The method for repairing the visually obscured area caused by suspended impurities for underwater structure detection according to claim 1 is characterized in that: The step S1 specifically includes the following steps: Step S11: Use Selflow optical flow method to obtain the current frame and adjacent frames and The motion optical flow field is converted into a visual color map of the adjacent frames. and ; Step S12: Using the Canny algorithm to visualize the color images of the adjacent frames obtained in step S11 and Perform edge detection, calculate the edge probability of each pixel, perform directional watershed transform and ultra-high-depth contour map transform (UCM) on the edge map, convert the open boundary probability area into several closed areas, and obtain the closed areas of the adjacent frames. and ; Step S13: The closed regions of the adjacent frames obtained in step S12 are respectively and The mean of the image segmentation threshold is used to obtain the visual color map of the adjacent frames before and after and The region segmentation result and ; Step S14: Visualize the optical flow color map of the adjacent frames obtained in step S11 and Convert to CIELAB color space, construct the color histogram in each area, and calculate the segmentation area Weighted contrast with other regions , get the foreground area probability estimation map of the adjacent frames before and after and : ; in, is the number of regions obtained by segmenting the image according to the image segmentation threshold; 、 They are all divided areas, and they are all and The area in ; and Indicates segmented area and The color histogram of , ; is the Euclidean distance, To segment the area and The distance from the center point of are the parameters of the spatial weighting scheme; Step S15: Use the maximum inter-class variance method to calculate the foreground region probability estimation map of the adjacent frames obtained in step S14 and The global image threshold is used to obtain the distinguishing mark map of the foreground and background areas of the adjacent frames. and ; Step S16: Compensate the foreground area according to the background displacement vector surrounding the foreground area.
3. The method for repairing the visually obscured area caused by suspended impurities for underwater structure detection according to claim 2 is characterized in that: The specific steps of step S16 are: Step S16a: transform the previous adjacent frame so that its background area is consistent with the background area of the current frame, and extract any foreground area , and perform expansion processing on it, and construct 8 connected domain windows with the edge pixels of the expansion area as the center in turn, calculate the mean of the optical flow vectors in all windows as the background displacement compensation value of the foreground area, and cyclically calculate the background displacement of all foreground areas as the compensation information of the foreground target area to obtain the background optical flow field after compensation of the previous adjacent frame , and then follow the same steps to obtain the background optical flow field after compensation of the adjacent frames : ; in, Represents the edge pixels of the foreground target area after expansion, Indicates the edge pixel of the previous adjacent frame The average optical flow vector value corresponding to the 8-connected domain, Indicates the edge pixel of the next adjacent frame The average optical flow vector value corresponding to the 8-connected domain, Compute the sign for the mean; Step S16b: Using the background optical flow field after compensation of the previous adjacent frame obtained in step S16a Background optical flow field after compensation with the adjacent frame Perform alignment transformation on the front and back adjacent frames respectively, that is, adjust the image coordinates accordingly according to the optical flow vector to obtain the aligned front adjacent frames Aligned with the next adjacent frame : ; in, Indicates the horizontal coordinate of the edge pixel point, Indicates the vertical coordinate of the edge pixel.
4. The method for repairing the visually obscured area caused by suspended impurities for underwater structure detection according to claim 1 is characterized in that: The step S2 specifically includes the following steps: Step 21: For the current frame and the previous adjacent frame aligned with the one obtained in step S1 and aligned adjacent frames Calculate the forward frame difference and the backward frame difference to obtain the difference map between the current frame and the previous adjacent frame The difference map between the current frame and the next adjacent frame , and use it as the motion information of dynamic perception: ; in, express and The pixel position in is the difference image threshold, which is used to filter the noise of the difference image in the three RGB channels respectively; Step 22: The difference image between the current frame and the previous adjacent frame obtained in step 21 The difference map between the current frame and the next adjacent frame Perform median filtering to obtain the forward motion feature map and backward motion feature map ; Step S23: Use the maximum inter-class variance method to obtain the forward motion feature map Threshold and backward motion feature map Threshold : ; in, represents the global image threshold obtained by the maximum inter-class variance method, Indicates the calculation of the mean, and the small-size suspended impurity detection image is obtained by full-value threshold segmentation and ; Step 24: Use Canny and UCM to get the current frame The region segmentation result ; Step 25: Extract the region segmentation results obtained in step 24 in sequence For each pixel point in the image, we calculate the mean of the forward motion feature value and the backward motion feature value in each area to generate a motion contrast perception heat map. and , and then obtain the refined landmark map of forward motion through threshold segmentation and the refined landmark map of backward motion ; Step 26: Compute the refined landmark map for forward motion and the refined landmark map of backward motion Motion contrast perception heat map obtained based on global optical flow field and The intersection of the two results yields the refinement of the large-scale suspended impurity region. and ,Right now , ; Step 27: Small size suspended impurity detection map and Refinement results of large-scale suspended impurity detection area and Combined with the above, the final detection results of suspended impurities are obtained. .
5. The method for repairing the visually obscured area caused by suspended impurities for underwater structure detection according to claim 1 is characterized in that: The step S3 specifically includes the following steps: Step 31: Build a hybrid restoration label map based on the best matching regions in adjacent frames. Indicates that, Represents a pixel, Indicates that the pixel is not blocked by suspended impurities and does not need to be repaired; when When , it means that the best matching point of the pixel is the previous adjacent frame The corresponding pixel point; when When , it means that the best matching point of the pixel is the next adjacent frame The corresponding pixel point; when When , it means that no valid matching point is found for the pixel in the three-frame segment for repair, and it is necessary to estimate it by learning the surrounding information; Step 32: Use the method of minimizing the label cost to solve the mixed repair label map , calculate the minimum label cost according to the matching conditions in step 31 : ; in, Used to measure the difference between the aligned adjacent frames and the current frame; Used to measure the smoothing cost between the best matching area in the aligned adjacent frames and the adjacent area in the current frame; is the regularization parameter, ; Indicates the aligned adjacent frames, including the aligned previous adjacent frames Aligned with the next adjacent frame ; It's a pixel The set of 4-connected neighbors; is the exclusive OR operator; and Respectively Pixels within the region; Step 33: By minimizing the label cost , get the mixed repair label map , and then set all the areas in the current frame that are not blocked by suspended impurities to 0.
6. The method for repairing the visually obscured area caused by suspended impurities for underwater structure detection according to claim 5 is characterized in that: The step 33 specifically includes the following steps: Step 33A: Repair the label map based on the mixture For the current frame Perform the first restoration and use the multi-band fusion method to eliminate the uneven illumination traces of the restored image. The specific process is as follows; Step 33A1: Use Gaussian pyramid to restore the current frame , aligned previous adjacent frames and the next adjacent frame Decomposed into multiple multi-band sub-images, As the bottom image of the Gaussian pyramid , and then through a Low-pass Gaussian kernel From the underlying image Start convolution and downsampling layer by layer to obtain other pyramid layer images, and finally build a Layer Gaussian pyramid image layer , Indicates that Downsampling times, Indicates the layer Gaussian pyramid image, Indicates the number of Gaussian pyramid image layers; Step 33A2: In order to reduce redundant information in the Gaussian pyramid, it is necessary to further construct a Laplacian pyramid. The Laplacian image layer is obtained by subtracting two adjacent layers of the Gaussian pyramid from bottom to top: ; in, Represents the Laplacian pyramid image layer; Select the mixed repair label The parts of are used as fusion masks respectively, and the combined pyramid is constructed: ; Step 33A3: Upsample the combined pyramid of each layer to the original image size and accumulate to obtain the first restoration result : ; Step 33B: For the areas in the adjacent frames that cannot provide effective repair information, the joint space-time transformation network STTN of video repair is used for further repair. The three frame segments are used as network input data, and the suspended impurity detection result area is As an occlusion mask, perform a second repair on the current frame to obtain the second repair result ; Step 33C: Target The structural information of the repaired area is blurred, resulting in the loss of internal details. and Perform fusion to obtain the final repaired image : ; in, is the set weight coefficient.
Citation Information
Patent Citations
A method for removing bubble noise of an underwater image
CN109166083A
Underwater image de-occlusion method based on generative adversarial network
CN111640075A