Interactive water body extraction method, system and equipment based on deep learning and medium
By pre-processing and interactive data encoding of the water body sample data set, combined with deep learning and fully connected conditional random field, efficient and automatic extraction of water body data in high-resolution remote sensing images is achieved, solving the problem of low water body extraction efficiency in the existing technology, and improving the accuracy of water body extraction and production application efficiency.
Patent Information
- Application Number
- CN202411984309.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-31
- Publication Date
- 2025-05-16
AI Technical Summary
The existing water extraction methods cannot automatically extract ground object targets in one go and high quality, and require manual verification and correction, which is inefficient and cannot meet the real-time data acquisition needs of high-resolution remote sensing images.
By obtaining the water body sample data set for median filtering and pre-processing, generating foreground and background sample line diagrams, simulating user interaction to generate interactive data encoding diagrams, using deep learning network to obtain water body image feature maps, and training and update through segmentation network model, combining with full-connection conditions to optimize water body boundaries, realizing automatic extraction of water body data.
It realizes high-quality and efficient water extraction, reduces the workload of manual correction, improves the accuracy and efficiency of water extraction, and is suitable for real-time data acquisition of high-resolution remote sensing images.
Smart Images

Figure CN120014406A_ABST
Abstract
Description
Technical Field
[0001] The invention belongs to the technical field of image processing and relates to a water body extraction method and system, equipment and medium. Background Art
[0002] Water bodies are an important part of natural resource monitoring. Water body data is resource information that ensures the normal operation of social and economic life. It is also an important object that needs to be paid attention to in projects such as engineering construction and operation and maintenance. For example, in the operation and maintenance of power grid projects, timely acquisition of accurate and effective water body data is conducive to taking measures to prevent water disasters and ensure the safety of power grid property.
[0003] Among the current water body extraction methods, the extraction method based on water body index is widely used in single-band and multi-band images, but it cannot completely suppress background information unrelated to water bodies, and the accuracy of the extraction results depends on expert knowledge, which limits its application scenarios; after the development of machine learning methods such as decision trees, random forests, and support vector machines, the water body extraction effect has been improved in some scenarios, but water body extraction is limited by factors such as expert knowledge and post-processing of extraction results, and the work efficiency is difficult to meet the needs of timely acquisition of water body data due to the increase in image data update speed. In recent years, deep learning technology based on convolutional neural networks has begun to be applied to remote sensing image water body extraction tasks, and feature extraction from massive labeled data has achieved good results. However, on the one hand, deep learning requires the input of a large amount of sample data, and the cost of manual collection by users is very high. On the other hand, whether it is a traditional model based on artificially designed features or a deep learning model based on powerful feature extraction capabilities, it is impossible to automatically extract ground objects from complex high-resolution remote sensing images at one time and with high quality. The prediction results still need to be manually checked and corrected, and the water body extraction efficiency is low, making production applications very inefficient. Summary of the invention
[0004] In order to solve the problems that the water body extraction method described in the background technology cannot automatically extract ground object targets in one time and with high quality, requires manual verification and correction of prediction results, is inefficient, and is very inefficient in production application, the present invention proposes an interactive water body extraction method, system, equipment and medium based on deep learning.
[0005] The method of the present invention comprises:
[0006] Obtain water sample dataset;
[0007] Perform median filtering preprocessing on the image data in the water body sample data set to obtain a preprocessed water body sample data set;
[0008] For the semantic label data set in the preprocessed water sample data set, respectively generating a foreground sample line graph and a background sample line graph based on a skeleton line extraction algorithm;
[0009] According to the foreground sample line graph and the background sample line graph, the points or lines are respectively used as positive and negative sample points or lines, and the points or lines are regarded as independent pixels or a set of pixels and marked as foreground or background, and the user interaction situation is simulated to generate an interaction data coding graph taking into account multiple features;
[0010] Based on the preprocessed water sample data set, the water image feature map is obtained through the deep learning network;
[0011] The image data in the preprocessed water body sample data set, the interactive data encoding map and the water body image feature map are used as the segmentation network model input, and the segmentation network model is trained and updated to obtain a water body interactive segmentation model;
[0012] Based on the water body interactive segmentation model, water body result data is extracted from high-resolution water body remote sensing image data to obtain a water body data probability map;
[0013] According to the water body data probability map, the water body boundary is optimized using a fully connected conditional random field, and a water body vector result in a vector format is output to achieve water body extraction.
[0014] Furthermore, the preprocessing method of the image data in the water sample data set is:
[0015] The median filter image balancing method is used to replace the value of the central pixel with the median of all pixels in the window to filter out isolated noise. The calculation formula is as follows:
[0016] y i,j =Median(x i,j ),x i,j ∈Α (1),
[0017] Among them, x i,j and i,j Represent the input and output image pixels respectively, Α represents the processing window of median filtering, and the Median function returns the median of the pixels in the window.
[0018] Furthermore, the method for generating the foreground sample line graph and the background sample line graph includes:
[0019] Firstly, the target object to be extracted is binarized according to the true image value obtained from the semantic label data set in the preprocessed water sample data set, and then the independent binary region of the current target is obtained by judging according to the four-connected regions;
[0020] Next, the skeleton line extraction algorithm is used. According to the independent binary region of the current target, the morphological thinning operation is used to erode the pixels at the boundary until it is impossible to thin it further, so that all pixels in the skeleton line belong to the inside of the target and do not contain any ambiguity, and the skeleton curve of the target with a single pixel width is obtained.
[0021] Finally, these single-pixel-wide skeleton lines are expanded according to the pixel binary labels to cover more pixels and generate foreground seed lines as foreground sample line maps;
[0022] After inverting the true value of the image obtained from the semantic label dataset, the background skeleton line is generated in the same way as the foreground seed line as above, as the background sample line map.
[0023] Furthermore, the interactive data encoding diagram includes: a correlation diagram E(v i,j ), binary coding graph BS(v i,j ) and positive and negative geodesic graphs GD p (v i,j ) and GD n (v i,j );
[0024] The correlation graph E(v i,j ) include:
[0025] For the semantic data set of the preprocessed water sample data set, the positive and negative sample points or lines obtained by simulating user interaction are generated according to the Euclidean distance transformation to generate a positive Euclidean distance map ED p , where p represents the positive channel graph, and the negative Euclidean distance graph ED can be obtained in the same way n , where n represents the negative channel map; the positive Euclidean distance map ED p , Negative Euclidean distance graph ED n Combine, integrate the pixel position into a channel map, recalculate the spatial position encoding value of each pixel in the image; normalize the spatial position encoding value of each pixel to [0,1], and obtain the correlation map E(v i,j ), E(v i,j ) is calculated as follows:
[0026]
[0027]
[0028]
[0029] Among them, v i,j represents the value of a pixel at any position (i, j) in the image, S is the set of labeled pixels, S p and S nRepresent the foreground labeled pixel set and the background labeled pixel set respectively;
[0030] The binary code map BS(v i,j ) is calculated as:
[0031] Following the user-labeled pixels, the original binary encoding map of the sampled pixels is retained, and the positive and negative sample points and lines are fused into a binary encoding map BS (v i,j ), set the value of the user-marked pixel to 255 and the values of the remaining pixels to 0;
[0032] The positive and negative geodesic graphs GD p (v i,j ) and GD n (v i,j ) is calculated as:
[0033] The geodesic distance transform is used to encode the user interaction using the rich information of the image itself, and the geodesic distance map GD is obtained. t (v i,j ), GD t (v i,j ) is calculated as follows:
[0034]
[0035]
[0036] Among them, I represents the entire input image, Path v,u and r(s) represent all paths and their direction vectors between pixel v and pixel u, respectively. Represents the gradient differential of the direction vector between two pixels; through calculation, the positive and negative geodesic maps GD can be calculated for the front and background marked pixels respectively. p (v i,j ) and GD n (v i,j ).
[0037] Furthermore, the method for obtaining the water body image feature map is:
[0038] The preprocessed water sample dataset is taken as input, and the parameters of the VGG network model pre-trained on the ImageNet dataset are fixed through a deep learning network to extract image features. The feature tensors of "conv1_2", "conv2_2", "conv3_2", "conv4_2", and "conv5_2" are sampled bilinearly to obtain a feature map with the same size as the original image as the water image feature map.
[0039] Furthermore, in the convolutional layer of the segmentation network model, 1×1 dilated convolution is used with an output channel of 64. In the downsampling part, cascaded dilated convolutions at a quarter of the image size are used to increase the dilated ratio to 1, 2, 4, 8, and 16 while ensuring that the number of output channels is always 64. Each dilated convolution is followed by a ReLU layer. Similarly, edge padding is used to keep the size of the tensor consistent.
[0040] In terms of the attention module of the segmentation network model, it is assumed that the input feature map is F = [F1, ..., F C ]∈R C ,H,W , where C represents the number of channels, H and W represent the height and width of the feature map respectively; first, two paths perform a global average pooling and a global maximum pooling operation on the input features, compressing the global spatial information and obtaining two results at the same time. gap ∈R C,1,1 and squeeze gmp ∈R C,1,1 ; Then use a multi-layer perceptron to stimulate the above results respectively, and add the results element by element, and then apply a sigmoid function and a scale transformation function to the stimulated results to obtain the weight map E weig ht ; Finally, the weight map E is multiplied pixel by pixel weig ht Assign weights to the input feature map and obtain the output of the attention module AGC , the calculation formula is as follows:
[0041]
[0042] in, represents the pixel-by-pixel multiplication operation, σ represents the sigmoid function and the scale function, and MLP represents a shared network consisting of a multilayer perceptron with one hidden layer;
[0043] During the training and updating process of the segmentation network model, the interactive data encoding map, the image data in the preprocessed water body sample data set, and the water body image feature map are used as input to predict the binary mask of the ground object target, and the target true value in the semantic label data set is used to update the network parameters to obtain the water body interactive segmentation model;
[0044] The loss function used in the training update process is defined as follows:
[0045] Loss = min loss δ (Y,P δ ) (9),
[0046]
[0047] Among them, Y and P δ They represent the true value of the label and the predicted result with parameter δ respectively, and v represents each pixel in the image.
[0048] Furthermore, the boundary optimization process includes:
[0049] A post-processing random field model is established using a fully connected conditional random field. A probability graph representing the probability value of each pixel being assigned to the foreground or background is obtained based on the water body prediction model. The water body remote sensing image data is sent to the post-processing random field model. The post-processing random field model contains a large number of nodes and edges. Each pixel in the water body remote sensing image data is a node in the graph model, and the line connecting adjacent pixels is an edge. Given a set of input random variables, the post-processing random field model outputs a conditional probability distribution model for another set of random variables. The formula is defined as follows:
[0050]
[0051] Among them, X, Y represent the input image and the corresponding binary image respectively, N represents the set of all pixels, and the threshold p of each pixel position i is {0,1}, the first part of the data term function φ i It describes the cost of assigning a pixel. The second part is the smoothing function φ ii It describes the cost of calculating the connectivity of similar pixels. The definitions of the two functions are as follows:
[0052]
[0053]
[0054] in, and They represent the probability that the network predicts that the pixel i belongs to the foreground and background, δ and k represent the penalty function and kernel function, respectively. i and f j represent the feature vectors of pixel points i and j in a certain space respectively; the penalty function δ limits the conduction of energy, if p i ≠p j ,δ(p i ,p j )=0; k represents a Gaussian kernel function, which is constrained by two weight parameters. The Gaussian kernel function used is defined as follows:
[0055]
[0056] Among them, w1 and w2 represent the two weight parameters of the constrained Gaussian kernel function respectively; the first part of the function depends on the coordinate difference of the pixel points and the spectral intensity difference, and the second part of the function only depends on the coordinate difference of the pixel points, c i and c j Respectively represent the coordinates of the i-th and j-th pixel points; I i and I j Represent the spectral intensity of the i-th and j-th pixel points respectively; parameter θ α ,θ β ,θ γ Constrain the weight relationship between these kernel function information.
[0057] The present invention also proposes an interactive water body extraction system based on deep learning, including a data acquisition module, a data preprocessing module, a sample line graph generation module, an interactive data coding graph generation module, an image feature graph acquisition module, a water body interactive segmentation model establishment module, a water body data probability graph acquisition module and a water body boundary optimization module.
[0058] The data acquisition module is used to acquire a water body sample data set.
[0059] The data preprocessing module is used to perform median filtering preprocessing on the image data in the water body sample data set to obtain a preprocessed water body sample data set.
[0060] The sample line graph generation module is used to generate a foreground sample line graph and a background sample line graph based on a skeleton line extraction algorithm for the semantic label data set in the preprocessed water body sample data set.
[0061] The interaction data coding map generation module is used to use the foreground sample line map and the background sample line map as positive and negative sample points or lines, respectively, regard the points or lines as independent pixels or a group of pixels and mark them as foreground or background, simulate the user interaction situation, and generate an interaction data coding map that takes into account multiple features.
[0062] The water body image feature map acquisition module is used to acquire the water body image feature map through a deep learning network based on the preprocessed water body sample data set.
[0063] The water body interaction segmentation model establishment module is used to use the image data in the preprocessed water body sample data set, the interaction data encoding map, and the water body image feature map as segmentation network model inputs, train and update the segmentation network model to obtain a water body interaction segmentation model.
[0064] The water body data probability map acquisition module is used to extract water body result data from high-resolution water body remote sensing image data based on a water body interactive segmentation model to obtain a water body data probability map.
[0065] The water body boundary optimization module is used to optimize the water body boundary based on the water body data probability map using a fully connected conditional random field, and output a water body vector result in a vector format to achieve water body extraction.
[0066] The present invention also proposes an electronic device, comprising: a memory and a processor, wherein the memory and the processor are communicatively connected to each other, the memory stores computer instructions, and the processor executes the computer instructions to implement the interactive water body extraction method based on deep learning as described above.
[0067] The present invention also proposes a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, it implements the interactive water body extraction method based on deep learning as described above.
[0068] Compared with the prior art, the present invention first pre-processes the water sample data set to generate a foreground sample line graph and a background sample line graph, and generates an interactive data coding map that takes into account multiple features by simulating user interaction. At the same time, the water image feature map is obtained through a deep learning network, and the image data, interactive data coding map, and water image feature map in the pre-processed water sample data set are used as input to train and update the segmentation network model to obtain a water body interactive segmentation model. According to the water body interactive segmentation model, a water body data probability map can be obtained, and the water body boundary is optimized using a fully connected conditional random field. The present invention uses the user's prior recognition ability to guide the water body interactive segmentation model to segment and extract water bodies, reduce non-water body data in the model prediction results, and use a fully connected conditional random field to optimize the water body boundary, try to obtain an accurate and smooth water body boundary, and reduce the workload of manually modifying the boundary. The present invention can automatically extract ground object targets at one time and with high quality, without the need for manual verification and correction of the prediction results, greatly improving the efficiency of water body extraction, and is accurate and efficient in production applications, and has strong practicality. BRIEF DESCRIPTION OF THE DRAWINGS
[0069] Figure 1 The figure is a flow chart of the method of the present invention.
[0070] Figure 2 A visualization diagram of the interaction data encoding graph.
[0071] Figure 3 Schematic diagram of the segmentation network model.
[0072] Figure 4 Schematic diagram of the multi-layer hole convolution combined with the segmentation network structure.
[0073] Figure 5 Schematic diagram of the attention module AGC.
[0074] Figure 6 This is a comparison chart of water extraction effects.
[0075] Figure 7 It is a system architecture diagram of the present invention. DETAILED DESCRIPTION
[0076] In order to make the technical problems, technical solutions and beneficial effects to be solved by the present application clearer, the present application is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not intended to limit the present application. The positional relationships described in the embodiments are consistent with those shown in the accompanying drawings.
[0077] Interactive water extraction method based on deep learning, the flowchart is as follows Figure 1 The details are as follows.
[0078] Get the water sample dataset.
[0079] Specifically, the water body sample data set can be obtained based on the geographical national conditions results data. Of course, the acquisition channel of the water body sample data set is not limited to the geographical national conditions results data.
[0080] The image data in the water body sample data set is preprocessed by median filtering to obtain the preprocessed water body sample data set.
[0081] High-quality data can provide more accurate information and reduce noise interference. Median filtering preprocessing of image data in water sample datasets can improve image quality, eliminate noise interference, and enhance the distinguishability of surface features.
[0082] Specifically, the median filter image balancing method is used to replace the value of the central pixel with the median of all pixels in the window to filter out isolated noise. The calculation formula is as follows:
[0083]
[0084] Among them, x i,j and i,j Represent the input and output image pixels respectively, Α represents the processing window of median filtering, and the Median function returns the median of the pixels in the window.
[0085] For the semantic label data set in the preprocessed water body sample data set, a foreground sample line graph and a background sample line graph are respectively generated based on a skeleton line extraction algorithm.
[0086] Specifically, the method for generating the foreground sample line graph and the background sample line graph is as follows:
[0087] Firstly, the target object to be extracted is binarized according to the true image value obtained from the semantic label data set in the preprocessed water sample data set, and then the independent binary region of the current target is obtained by judging according to the four-connected regions;
[0088] Next, the skeleton line extraction algorithm is used. According to the independent binary region of the current target, the morphological thinning operation is used to erode the pixels at the boundary until it is impossible to thin it further, so that all pixels in the skeleton line belong to the inside of the target and do not contain any ambiguity, and the skeleton curve of the target with a single pixel width is obtained.
[0089] Finally, these single-pixel-wide skeleton lines are expanded according to the pixel binary labels to cover more pixels and generate foreground seed lines as foreground sample line maps;
[0090] After inverting the true value of the image obtained from the semantic label dataset, the background skeleton line is generated in the same way as the foreground seed line as above, as the background sample line map.
[0091] Since the background area does not belong to the area of interest of the extraction algorithm, there is no need to strictly limit the position and number of seed lines. In order to make the seed line more representative, the entire skeleton line can be cropped to generate shorter sub-lines to represent the background area.
[0092] According to the foreground sample line graph and the background sample line graph, the points or lines are respectively used as positive and negative sample points or lines, and the points or lines are regarded as independent pixels or a set of pixels and marked as foreground or background, and the user interaction is simulated to generate an interaction data encoding map taking into account multiple features. Among them, the foreground can also be called the "selected object" and the background can also be called the "area to be eliminated".
[0093] Specifically, the interaction data encoding graph is: i,j ), binary coding graph BS(v i,j ) and positive and negative geodesic graphs GD p (v i,j ) and GD n (v i,j ).
[0094] Correlation graph E(v i,j ) include:
[0095] For the semantic data set of the preprocessed water sample data set, the positive and negative sample points or lines obtained by simulating user interaction are generated according to the Euclidean distance transformation to generate a positive Euclidean distance map ED p , where p represents the positive channel graph, and the negative Euclidean distance graph ED can be obtained in the same way n , where n represents the negative channel map; the positive Euclidean distance map ED p , Negative Euclidean distance graph EDn Combine, integrate the pixel position into a channel map, recalculate the spatial position encoding value of each pixel in the image; normalize the spatial position encoding value of each pixel to [0,1], and obtain the correlation map E(v i,j ), the purpose of normalization is to more conveniently combine the correlation graph with other feature tensors to facilitate the gradient conduction of the model, E(v i,j ) is calculated as follows:
[0096]
[0097]
[0098]
[0099] Among them, v i,j represents the value of a pixel at any position (i, j) in the image, S is the set of labeled pixels, S p and S n They represent the foreground labeled pixel set and the background labeled pixel set respectively.
[0100] In addition to the spatial position relationship of pixels, geodesic distance transform can also be used to better utilize the rich information of the image itself to encode user interactions.
[0101] Positive and negative geodesic diagram GD p (v i,j ) and GD n (v i,j ) is calculated as:
[0102] The geodesic distance transform is used to encode the user interaction using the rich information of the image itself, and the geodesic distance map GD is obtained. t (v i,j ), GD t (v i,j ) is calculated as follows:
[0103]
[0104] Among them, I represents the entire input image, Path v,u and r(s) represent all paths and their direction vectors between pixel v and pixel u, respectively. Represents the gradient differential of the direction vector between two pixels; through calculation, the positive and negative geodesic maps GD can be calculated for the front and background marked pixels respectively. p (v i,j ) and GD n (v i,j ).
[0105] Binary coded graph BS(v i,j ) is calculated as:
[0106] Following the user-labeled pixels, the original binary encoding map of the sampled pixels is retained, and the positive and negative sample points and lines are fused into a binary encoding map BS (v i,j ), the value of the user-marked pixel is set to 255, and the values of the remaining pixels are set to 0.
[0107] The visualization diagram of the interactive data encoding diagram is as follows: Figure 2 As shown, (a) represents the original image with interaction, the red marks are negative sample points, the green marks are positive sample points, (b) represents the binary encoding map, (c) represents the Euclidean distance encoding map of the positive sample points, (d) represents the geodesic distance map of the positive sample points, and (e) represents the correlation map.
[0108] Based on the preprocessed water sample dataset, the water body image feature map is obtained through the deep learning network.
[0109] Specifically, the preprocessed water sample dataset is taken as input, and the parameters of the VGG network model pre-trained on the ImageNet dataset are fixed through a deep learning network to extract image features. The feature tensors of "conv1_2", "conv2_2", "conv3_2", "conv4_2", and "conv5_2" are sampled bilinearly to obtain a feature map with the same size as the original image as the water image feature map.
[0110] The image data, interactive data encoding map and water body image feature map in the preprocessed water body sample data set are used as the input of the segmentation network model, and the segmentation network model is trained and updated to obtain the water body interactive segmentation model.
[0111] Specifically, the schematic diagram of the segmentation network model is as follows Figure 3 As shown in the figure, the segmentation network design is divided into three parts: downsampling part, multi-layer hole convolution, and upsampling part.
[0112] Specifically, in the convolutional layer of the segmentation network model, the multi-layer hole convolution combined with the segmentation network structure is used. Figure 4As shown in the figure, since the number of input channels is too large, 1×1 dilated convolution (output channels are 64) is first used to reduce the number of channels in the shallow network. On the one hand, the dimensionality reduction method is used to reduce computing resources and facilitate the processing of large feature data; on the other hand, these input channels are not from the same source and should not be treated in the same way. They need to be processed initially to distinguish them. In the downsampling part, cascaded dilated convolutions at a quarter of the image size are used. While ensuring that the number of output channels is always 64, the dilation rate is expanded to 1, 2, 4, 8, and 16. Each dilated convolution is followed by a ReLU layer; similarly, edge padding is used to keep the size of the tensor consistent.
[0113] In terms of the attention module of the segmentation network model, the attention module AGC used is as follows Figure 5 As shown, assuming that the input feature map is F = [F1, ..., F C ]∈R C,H,W , where C represents the number of channels, H and W represent the height and width of the feature map respectively; first, two paths perform a global average pooling and a global maximum pooling operation on the input features, respectively, compressing the global spatial information while obtaining two results squeeze gap ∈R C,1,1 and squeeze gmp ∈R C,1,1 ; Then use a multi-layer perceptron to stimulate the above results respectively, and add the results element by element, and then apply a sigmoid function and a scale transformation function to the stimulated results to obtain the weight map E weight ; Finally, the weight map E is multiplied pixel by pixel weig ht Assign weights to the input feature map and obtain the output of the attention module AGC , the calculation formula is as follows:
[0114]
[0115] E weig ht =σ[MLP(squeeze gap )+MLP(squeeze gmp )] (8),
[0116] in, represents the pixel-by-pixel multiplication operation, σ represents the sigmoid function and the scale function, and MLP represents a shared network consisting of a multilayer perceptron with one hidden layer;
[0117] In the training and updating process of the segmentation network model, the interactive data encoding map, the image data in the preprocessed water sample data set, and the water image feature map are used as input to predict the binary mask of the ground object target, and the target true value in the semantic label data set is used to update the network parameters to obtain the water body interactive segmentation model;
[0118] The loss function used in the training update process is defined as follows:
[0119] Loss = minloss δ (Y,P δ ) (9),
[0120]
[0121] Among them, Y and P δ They represent the true value of the label and the predicted result with parameter δ respectively, and v represents each pixel in the image.
[0122] Based on the water body interactive segmentation model, water body result data is extracted from high-resolution water body remote sensing image data to obtain a water body data probability map.
[0123] According to the water body data probability map, the water body boundary is optimized using a fully connected conditional random field, and a water body vector result in a vector format is output to achieve water body extraction.
[0124] Specifically, the boundary optimization process includes:
[0125] A post-processing random field model is established using a fully connected conditional random field. A probability graph representing the probability value of each pixel being assigned to the foreground or background is obtained based on the water body prediction model. The water body remote sensing image data is sent to the post-processing random field model. The post-processing random field model contains a large number of nodes and edges. Each pixel in the water body remote sensing image data is a node in the graph model, and the line connecting adjacent pixels is an edge. Given a set of input random variables, the post-processing random field model outputs a conditional probability distribution model for another set of random variables. The formula is defined as follows:
[0126]
[0127] Among them, X and Y represent the input image and the corresponding binary image respectively, N represents the set of all pixels, and the threshold p of each pixel position i is {0,1}, the first part of the data term function φ i It describes the cost of assigning a pixel. The second part is the smoothing function φ i,j It describes the cost of calculating the connectivity of similar pixels. The definitions of the two functions are as follows:
[0128]
[0129]
[0130] in, and They represent the probability that the network predicts that the pixel i belongs to the foreground and background, δ and k represent the penalty function and kernel function, respectively. i and f j represent the feature vectors of pixel points i and j in a certain space respectively; the penalty function δ limits the conduction of energy, if p i ≠p j ,δ(p i ,p j )=0, indicating that if they are not equal, the function value is zero. The purpose is to conduct energy only when the labels between pixels are the same; k represents a Gaussian kernel function, which is constrained by two weight parameters. The Gaussian kernel function used is defined as follows:
[0131]
[0132] Among them, w1 and w2 represent the two weight parameters of the constrained Gaussian kernel function respectively; the first part of the function depends on the coordinate difference of the pixel points and the spectral intensity difference, and the second part of the function only depends on the coordinate difference of the pixel points, c i and c j Respectively represent the coordinates of the i-th and j-th pixel points; I i and I j Represent the spectral intensity of the i-th and j-th pixel points respectively; parameter θ α ,θ β ,θ γ Constrain the weight relationship between these kernel function information.
[0133] In this embodiment, based on experience and experimental results, these parameters are set as follows: w1, w2, θ α ,θ β ,θ γ 8, 10, 40, 18, 3, Figure 6 A comparison chart of water extraction effects is shown, Figure 6 (a) is the water body vectorization result directly based on the prediction result. Figure 6 (b) is the water body extraction result after boundary optimization. It can be clearly seen that after boundary optimization, the water body boundary is more accurate and smoother.
[0134] Interactive water extraction system based on deep learning, the architecture diagram is as follows Figure 7As shown, it consists of a data acquisition module, a data preprocessing module, a sample line graph generation module, an interactive data encoding graph generation module, an image feature graph acquisition module, a water body interactive segmentation model establishment module, a water body data probability graph acquisition module and a water body boundary optimization module.
[0135] The data acquisition module is used to obtain water sample data sets.
[0136] The data preprocessing module is used to perform median filtering preprocessing on the image data in the water body sample data set to obtain the preprocessed water body sample data set.
[0137] The sample line graph generation module is used to generate a foreground sample line graph and a background sample line graph based on a skeleton line extraction algorithm for the semantic label data set in the preprocessed water body sample data set.
[0138] The interactive data coding map generation module is used to use the foreground sample line map and the background sample line map as positive and negative sample points or lines, respectively, regard the points or lines as independent pixels or a group of pixels and mark them as foreground or background, simulate the user interaction situation, and generate an interactive data coding map that takes into account multiple features.
[0139] The water body image feature map acquisition module is used to obtain the water body image feature map through a deep learning network based on the preprocessed water body sample data set.
[0140] The water body interaction segmentation model establishment module is used to use the image data in the preprocessed water body sample data set, the interaction data encoding map, and the water body image feature map as segmentation network model inputs, train and update the segmentation network model to obtain a water body interaction segmentation model.
[0141] The water body data probability map acquisition module is used to extract water body result data from high-resolution water body remote sensing image data based on the water body interactive segmentation model to obtain a water body data probability map.
[0142] The water body boundary optimization module is used to optimize the water body boundary based on the water body data probability map using a fully connected conditional random field, and output the water body vector result in a vector format to achieve water body extraction.
[0143] The specific implementation method of each module in the system is consistent with that described in the above method and will not be repeated here.
[0144] The present invention also proposes an electronic device, comprising: a memory and a processor, the memory and the processor are communicatively connected to each other, the memory stores computer instructions, and the processor executes the computer instructions to implement the interactive water body extraction method based on deep learning and the interactive water body extraction system based on deep learning as described above.
[0145] The present invention also proposes a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, it implements the interactive water body extraction method based on deep learning and the interactive water body extraction system based on deep learning as described above.
[0146] Those skilled in the art will appreciate that the embodiments of the present application can be provided as methods, systems, or computer program products. Therefore, the application can adopt the form of complete hardware embodiment, complete software embodiment, or the embodiment in combination with software and hardware. Moreover, the application can adopt the form of the computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) that contain computer-usable program code. The scheme in the embodiment of the present application can be implemented in various computer languages, for example, object-oriented programming language Java, C++, Python and literal scripting language JavaScript, etc.
[0147] The present application is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of the processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0148] These computer program instructions may also be stored in a computer-readable memory capable of directing a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.
[0149] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for implementing the process in the computer or other programmable device. Figure 1 A process or multiple processes and / or boxes Figure 1 The steps for the functions specified in one or more boxes.
[0150] Although the preferred embodiments of the present application have been described, those skilled in the art may make other changes and modifications to these embodiments once they have learned the basic creative concept. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments and all changes and modifications falling within the scope of the present application.
[0151] Obviously, those skilled in the art can make various changes and modifications to the present application without departing from the spirit and scope of the present application. Thus, if these modifications and variations of the present application fall within the scope of the claims of the present application and their equivalents, the present application is also intended to include these modifications and variations.
Claims
1. An interactive water extraction method based on deep learning, characterized in that: include: Obtain water sample dataset; Perform median filtering preprocessing on the image data in the water body sample data set to obtain a preprocessed water body sample data set; For the semantic label data set in the preprocessed water sample data set, respectively generating a foreground sample line graph and a background sample line graph based on a skeleton line extraction algorithm; According to the foreground sample line graph and the background sample line graph, the points or lines are respectively used as positive and negative sample points or lines, and the points or lines are regarded as independent pixels or a set of pixels and marked as foreground or background, and the user interaction situation is simulated to generate an interaction data coding graph taking into account multiple features; Based on the preprocessed water sample data set, the water image feature map is obtained through the deep learning network; The image data in the preprocessed water body sample data set, the interactive data encoding map and the water body image feature map are used as the segmentation network model input, and the segmentation network model is trained and updated to obtain a water body interactive segmentation model; Based on the water body interactive segmentation model, water body result data is extracted from high-resolution water body remote sensing image data to obtain a water body data probability map; According to the water body data probability map, the water body boundary is optimized using a fully connected conditional random field, and a water body vector result in a vector format is output to achieve water body extraction.
2. The interactive water body extraction method based on deep learning according to claim 1, characterized in that: The preprocessing method of the image data in the water sample dataset is: The median filter image balancing method is used to replace the value of the central pixel with the median of all pixels in the window to filter out isolated noise. The calculation formula is as follows: y i,j =Median(x i,j 0,x i,j ∈Α (1), Among them, x i,j and i,j Represent the input and output image pixels respectively, Α represents the processing window of median filtering, and the Median function returns the median of the pixels in the window.
3. The interactive water body extraction method based on deep learning according to claim 2, characterized in that: The method for generating the foreground sample line graph and the background sample line graph comprises: Firstly, the target object to be extracted is binarized according to the true image value obtained from the semantic label data set in the preprocessed water sample data set, and then the independent binary region of the current target is obtained by judging according to the four-connected regions; Next, the skeleton line extraction algorithm is used. According to the independent binary region of the current target, the morphological thinning operation is used to erode the pixels at the boundary until it is impossible to thin it further, so that all pixels in the skeleton line belong to the inside of the target and do not contain any ambiguity, and the skeleton curve of the target with a single pixel width is obtained. Finally, these single-pixel-wide skeleton lines are expanded according to the pixel binary labels to cover more pixels and generate foreground seed lines as foreground sample line maps; After inverting the true value of the image obtained from the semantic label dataset, the background skeleton line is generated in the same way as the foreground seed line as above, as the background sample line map.
4. The interactive water body extraction method based on deep learning according to claim 3, characterized in that: The interactive data coding diagram includes: i,j ), binary coding graph BS(v i,j ) and positive and negative geodesic graphs GD p (v i,j ) and GD n (V i,j ); The correlation graph E(v i,j ) include: For the semantic data set of the preprocessed water sample data set, the positive and negative sample points or lines obtained by simulating user interaction are generated according to the Euclidean distance transformation to generate a positive Euclidean distance map ED p , where p represents the positive channel graph, and the negative Euclidean distance graph ED can be obtained in the same way n , where n represents the negative channel map; the positive Euclidean distance map ED p , Negative Euclidean distance graph ED n Combine, integrate the pixel position into a channel map, recalculate the spatial position encoding value of each pixel in the image; normalize the spatial position encoding value of each pixel to [0, 1], and obtain the correlation map E(v i,j ), E(v i,j ) is calculated as follows: Among them, v i,j represents the value of a pixel at any position (i, j) in the image, S is the set of labeled pixels, S p and S n Represent the foreground labeled pixel set and the background labeled pixel set respectively; The binary code map BS(v i,j The calculation method of 0 is: Following the user-labeled pixels, the original binary encoding map of the sampled pixels is retained, and the positive and negative sample points and lines are fused into a binary encoding map BS (v i,j ), set the value of the user-marked pixel to 255 and the values of the remaining pixels to 0; The positive and negative geodesic graphs GD p (v i,j ) and GD n (v i,j ) is calculated as: The geodesic distance transform is used to encode the user interaction using the rich information of the image itself, and the geodesic distance map GD is obtained. t (v i,j ), GD t (v i,j ) is calculated as follows: Among them, I represents the entire input image, Path v,u and r(s) represent all paths and their direction vectors between pixel v and pixel u, respectively. Represents the gradient differential of the direction vector between two pixels; through calculation, the positive and negative geodesic maps GD can be calculated for the front and background marked pixels respectively. p (v i,j ) and GD n (v i,j ).
5. The interactive water body extraction method based on deep learning according to claim 4, characterized in that: The method for obtaining the water body image feature map is: The preprocessed water sample dataset is used as input. Through the deep learning network, the parameters of the VGG network model pre-trained on the ImageNet dataset are fixed to extract image features. The feature tensors of "conv1_2", "conv2_2", "conv3_2", "conv4_2", and "conv5_2" are bilinearly sampled to obtain a feature map with the same size as the original image as the image feature map.
6. The interactive water body extraction method based on deep learning according to claim 5, characterized in that: In the convolutional layer of the segmentation network model, 1×1 dilated convolution is used with 64 output channels. In the downsampling part, cascaded dilated convolutions at a quarter of the image size are used to ensure that the number of output channels is always 64, while expanding the dilated ratio to 1, 2, 4, 8, and 16. Each dilated convolution is followed by a ReLU layer. Similarly, edge padding is used to keep the size of the tensor consistent. In terms of the attention module of the segmentation network model, it is assumed that the input feature map is F = [F1, ..., F c ]∈R c,H,W , where C represents the number of channels, H and W represent the height and width of the feature map respectively; first, two paths perform a global average pooling and a global maximum pooling operation on the input features, respectively, compressing the global spatial information while obtaining two results squeeze gap ∈R C,1,1 and squeeze gmp ∈R C,1,1 ; Then use a multi-layer perceptron to stimulate the above results respectively, and add the results element by element, and then apply a sigmoid function and a scale transformation function to the stimulated results to obtain the weight map E weight ; Finally, the weight map E is multiplied pixel by pixel weight Assign weights to the input feature map and obtain the output of the attention module AGC , the calculation formula is as follows: E weight =σ[MLP(squeeze gap )+MLP(squeeze gmp )] (8), in, represents the pixel-by-pixel multiplication operation, σ represents the sigmoid function and the scale function, and MLP represents a shared network consisting of a multilayer perceptron with one hidden layer; During the training and updating process of the segmentation network model, the interactive data encoding map, the image data in the preprocessed water body sample data set, and the water body image feature map are used as input to predict the binary mask of the ground object target, and the target true value in the semantic label data set is used to update the network parameters to obtain the water body interactive segmentation model; The loss function used in the training update process is defined as follows: Among them, Y and P δ They represent the true value of the label and the predicted result with parameter δ respectively, and v represents each pixel in the image.
7. The interactive water body extraction method based on deep learning according to claim 6, characterized in that: The boundary optimization process includes: A post-processing random field model is established using a fully connected conditional random field. A probability graph representing the probability value of each pixel being assigned to the foreground or background is obtained based on the water body prediction model. The water body remote sensing image data is sent to the post-processing random field model. The post-processing random field model contains a large number of nodes and edges. Each pixel in the water body remote sensing image data is a node in the graph model, and the line connecting adjacent pixels is an edge. Given a set of input random variables, the post-processing random field model outputs a conditional probability distribution model for another set of random variables. The formula is defined as follows: Among them, X and Y represent the input image and the corresponding binary image respectively, N represents the set of all pixels, and the threshold p of each pixel position i is {0,1}, the first part of the data item function Describes the cost of assigning a pixel. The second part is the smoothing function It describes the calculation cost of maintaining the connectivity of similar pixels. The definitions of the two functions are as follows: The definitions of the two functions are as follows: in, and They represent the probability that the network predicts that the pixel i belongs to the foreground and background, δ and k represent the penalty function and kernel function, respectively. i and f j represent the feature vectors of pixel points i and j in a certain space respectively; the penalty function δ limits the conduction of energy, if p i ≠p j ,δ(p i ,p j )=0; k represents a Gaussian kernel function, which is constrained by two weight parameters. The Gaussian kernel function used is defined as follows: Among them, w1 and w2 represent the two weight parameters of the constrained Gaussian kernel function respectively; the first part of the function depends on the coordinate difference of the pixel points and the spectral intensity difference, and the second part of the function only depends on the coordinate difference of the pixel points, c i and c j Respectively represent the coordinates of the i-th and j-th pixel points; I i and I j Represent the spectral intensity of the i-th and j-th pixel points respectively; parameter θ α ,θ β ,θ γ Constrain the weight relationship between these kernel function information.
8. Interactive water extraction system based on deep learning, characterized by: It includes a data acquisition module, a data preprocessing module, a sample line map generation module, an interactive data coding map generation module, a water body image feature map acquisition module, a water body interactive segmentation model establishment module, a water body data probability map acquisition module and a water body boundary optimization module; The data acquisition module is used to acquire a water sample data set; The data preprocessing module is used to perform median filtering preprocessing on the image data in the water body sample data set to obtain a preprocessed water body sample data set; The sample line graph generation module is used to generate a foreground sample line graph and a background sample line graph based on a skeleton line extraction algorithm for the semantic label data set in the preprocessed water body sample data set; The interaction data coding map generation module is used to use the foreground sample line map and the background sample line map as positive and negative sample points or lines, respectively, regard the points or lines as independent pixels or a set of pixels and mark them as foreground or background, simulate the user interaction situation, and generate an interaction data coding map taking into account multiple features; The water body image feature map acquisition module is used to acquire the water body image feature map through a deep learning network based on the preprocessed water body sample data set; The water body interactive segmentation model establishment module is used to use the image data in the preprocessed water body sample data set, the interactive data encoding map and the water body image feature map as the segmentation network model input, train and update the segmentation network model to obtain the water body interactive segmentation model; The water body data probability map acquisition module is used to extract water body result data from high-resolution water body remote sensing image data according to the water body interactive segmentation model to obtain a water body data probability map; The water body boundary optimization module is used to optimize the water body boundary based on the water body data probability map using a fully connected conditional random field, and output a water body vector result in a vector format to achieve water body extraction.
9. An electronic device, characterized in that: include: A memory and a processor, wherein the memory and the processor are communicatively connected to each other, the memory stores computer instructions, and the processor implements the interactive water body extraction method based on deep learning as described in any one of claims 1 to 7 by executing the computer instructions.
10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the interactive water body extraction method based on deep learning as described in any one of claims 1 to 7 is implemented.