A method for removing water surface reflections and a method for detecting navigable areas for unmanned boats
Image style conversion through the RDGAN network removes or weakens the reflection of water surface and highlight reflection, solving the problem of low detection accuracy of navigable water surface areas in complex water surface environments, and achieving higher detection accuracy and image clarity.
Patent Information
- Application Number
- CN202410139913.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-01-31
- Publication Date
- 2025-05-09
- Estimated Expiration
- 2044-01-31
AI Technical Summary
The prior art is difficult to effectively remove water surface reflections and highlight reflections in complex water surface environments, resulting in low detection accuracy of navigable areas of water surface.
The image style conversion is adopted using the RDGAN network, and the clarity of the water surface image is improved through the improved CycleGAN model and AdaIN encoder.
It significantly improves the accuracy of navigable area detection on the water surface, solves the semantic pixel misclassification problem caused by water surface interference, and the generated image is more realistic and the background remains unchanged.
Smart Images

Figure CN118155163B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of image noise reduction and intelligent perception of unmanned ships, and specifically to a method for removing water surface reflections and a method for detecting navigable areas on the water surface of unmanned ships. Background Art
[0002] In recent years, unmanned watercraft (USV) technology has played an increasingly important role in patrolling, monitoring, security and other tasks. Surface navigable area detection plays an important role in many applications such as unmanned vessel perception systems and water mapping. By estimating the navigable area on the water surface through the water body image segmentation method, it can provide a safe driving range for surface automatic navigation and provide an important reference for USV path planning, obstacle avoidance and target detection.
[0003] However, detecting navigable areas on the water surface is a challenging task. Existing cameras will have shadows on cloudy days or at night and in low-light environments, resulting in unclear images. In addition, the water surface will reflect strong light under the sunlight on a sunny day. Even if the water surface reflects sunlight well on a sunny day, reflections of river banks, buildings, vegetation, etc. will appear on the water surface, which will confuse the semantic feature information of the acquired river surface image. For segmentation tasks, these are all noises, and common water surface segmentation algorithms cannot work properly. In short, in practical applications, there is a lot of noise interference in the images directly obtained by USV, especially in the following situations:
[0004] (1) There are reflections of vegetation and buildings on the shore on the water surface. Conventional image semantic segmentation methods are difficult to determine the water surface area based only on texture features.
[0005] (2) When the sunlight is strong, there are many rays of light and reflections on the water surface. When the sunlight is incident on the water surface, it forms strong reflected radiation in the mirror direction, forming the mirror reflection of the water body.
[0006] (3) There are buildings blocking the edge of the waterside. Objects will be shadowed when there is insufficient light. There are sudden changes in light inside and outside the bridge. The boundary between the riverbank and the water surface is unclear, making it difficult to determine the location of the waterside boundary line.
[0007] All of the above conditions will make it difficult for image processing algorithms to perceive the water surface environment. Therefore, there is an urgent need for a method that can solve the problem of complex water surface segmentation. It is necessary to improve the accuracy of detecting navigable areas on the water surface in an environment with strong reflections and strong high-light reflections on the water surface. Summary of the invention
[0008] The object of the present invention is to provide a method for removing water surface reflections and reflections and a method for detecting navigable areas on the water surface for unmanned boats, so as to solve the problems raised in the above-mentioned background technology.
[0009] To achieve the above object, the present invention provides the following technical solutions:
[0010] A method for removing water surface reflections and highlights, which removes or weakens water surface reflections and highlight reflections by using an image style conversion method;
[0011] The image style transfer method adopts the RDGAN network, and the construction of the RDGAN network includes the following steps:
[0012] Step 1: X and Y are high-noise style source domain images and low-noise style target domain images, respectively. The two generators G in the original cycle consistency generative adversarial network CycleGAN model are X→Y and G Y→X Improved to use only a single generator G with a lightweight AdaIN encoder (AdaINC) implementation;
[0013] Step 2: In the AdaIN encoder, the adaptive instance normalization (AdaIN) module and the spatial adaptive normalization (SPADE) module are alternately used for normalization;
[0014] In the decoder, the adaptive instance normalization (AdaIN) module and the content and spatial adaptive normalization (CSPADE) module are used alternately for normalization;
[0015] Step 3, controlling the quality of the generated water surface denoised image by generating AdaIN harmonic interpolation;
[0016] Step 4: Modify the objective loss function of training optimization in the CycleGAN model and add feature matching loss and cycle semantic consistency loss.
[0017] As a further solution of the present invention: AdaIN in the AdaIN module adaptively generates affine parameters according to the input style encoding, and the normalized calculation formula is as follows:
[0018]
[0019] where μ i (y) and σ i (y) is the mean and variance of the target style, each feature map x i Use the affine parameters corresponding to style y to perform scaling and offset to achieve normalization;
[0020] AdaIN aligns the normalized channel mean and variance of images containing high noise content of water surfaces with the mean and variance of weak noise pattern images so that the generated images have the same feature distribution as low-noise water areas.
[0021] As a further solution of the present invention: the AdaIN affine parameter K of the AdaIN module is obtained by the AdaIN encoder and is defined as follows:
[0022]
[0023] Where F is the AdaIN code generator, v is the input random Gaussian vector, and F m The mapping network consists of MLP and fully connected layers, X and Y are high-noise style source domain images and low-noise style target domain images respectively; μ i (y) and σ i (y) is the mean and variance of the low-noise water surface image in the target domain, and is the normalized affine transformation parameter of AdaIN;
[0024] During the forward conversion in the training phase, the AdaIN encoder that converts the high-noise water surface image X into the low-noise water surface image Y is set to a fixed constant vector, specifically composed of a vector with a mean of 0 and a variance of 1, F(v) = K 0 =(0,1);
[0025] In the reverse conversion, the AdaIN encoding of the low-noise water surface image Y into the high-noise water surface image X is generated by the mapping network F m The generated vector is obtained.
[0026] As a further solution of the present invention: the quality of the generated image is controlled by interpolating the generated AdaIN code, and the AdaIN code interpolation calculation method is:
[0027] I(θ)=(1-θ)K 0 +θF m (v),0≤θ≤1
[0028] where I(θ) is between K 0 and F m (v) AdaIN harmonic interpolation, θ is the AdaIN harmonic interpolation coefficient;
[0029] Larger values of θ indicate more transition to high-noise features, while smaller values preserve more of low-noise features.
[0030] As a further solution of the present invention: the two generators in step 1 are improved to use only a single generator G, which is implemented by AdaIN encoder. The reversible conversion generator from domain X to domain Y is defined as
[0031] G X→Y =G(x,I(0))
[0032] G Y→X =G(y,I(1))
[0033] Only one generator G(x) and AdaIN harmonic interpolation I(.) are used to achieve the conversion between the two domains;
[0034] I(.) is the AdaIN harmonic interpolation.
[0035] As a further solution of the present invention: the CSPADE module in step 1 is obtained by fusing content feature information in the SPADE module as input;
[0036] Affine parameters and It is the joint input semantic segmentation mask m at the i-th layer feature map x and content encoding c i Obtained, n∈N,c∈C i ,y∈H i ,x∈W i ;
[0037] The specific definition of affine transformation normalization is:
[0038]
[0039] in, is the activation value of the i-th layer, the scaled value and deviation value They are the semantic masks m x and content information c i The weighted sum of is calculated as:
[0040]
[0041]
[0042] Among them, m x is the semantic mask information of the corresponding layer after resizing, c i It is the content feature encoding of the corresponding scale in the low-dimensional space. and represents the mean and standard deviation and are defined as follows:
[0043]
[0044] Where N is the number of samples, H i and W i are the height and width of layer i respectively; when the weight ω γ = 0, CSPADE will degenerate into the original SPADE, so it is set to ω in the encoder network of the generator during training. γ =0, and in the high-dimensional space of the decoding network, set ω γ>0, obtain the adaptability of content and space in high-level space, and this parameter value is obtained through training;
[0045] In the training stage, the semantic mask is provided by the semantic labels annotated in the training set images, while in the inference stage, the semantic mask is given an initial input as an estimate by the pre-trained lightweight FCN network; then the semantic information is corrected and calibrated through the RDGAN network.
[0046] As a further solution of the present invention: the residual block of AdaIN includes convolution, AdaIN, LeakyRelu and skip connection;
[0047] The AdaIN residual module consists of the AdaIN encoded features f i As an additional input, f i The feature map is normalized by adjusting the output mean and variance after the activation value as the affine transformation parameters of the AdaIN layer;
[0048] The residual block SPADEResBlock of SPADE includes convolution, AdaIN, LeakyRelu and skip connection, and the additional input of the modulation parameter is the semantic mask m of different scales x ;
[0049] The residual block of the CSPADE includes convolution, AdaIN, LeakyRelu and skip connection, and the additional input of the modulation parameters includes semantic masks m of different scales. x and content feature encoding at different scales c i .
[0050] As a further solution of the present invention: the AdaINC is a mapping network composed of multiple multi-layer perceptrons MLP and fully connected layers, the input of the mapping network is a randomly sampled 1*64 10-bit Gaussian noise vector v; the mapping network learns the noise-free or weak noise style features of the target domain from the vector v through MLP; the mapping network includes multiple fully connected layers, the last layer of the mapping network is not shared, and the remaining layers share parameters; in the last layer, linear activation is used for the generation of the affine parameter mean vector, and ReLU activation function is used for the generation of variance to ensure non-negativity.
[0051] As a further solution of the present invention: the stability of the training model is improved by improving the loss function of the CycleGAN model, the loss function includes feature matching loss and cycle semantic consistency loss
[0052] The loss function of the CycleGAN network training objective is defined as:
[0053]
[0054] in, is a confrontational loss, is the cycle consistency loss, It is a loss of identity. is the feature matching loss, is the cycle semantic consistency loss, λ cyc ,λ ident ,λ fm ,λ sem Respectively represent the weight coefficients of the corresponding losses;
[0055] Training is done by solving the following min-max objective function:
[0056]
[0057] Adversarial Loss:
[0058] The adversarial loss with a single generator is expressed as follows:
[0059]
[0060] in‖·‖ 2 is the L2 norm, x∈X, x∈Y, G(x) represents the generated low-noise image, D X and D Y is the discriminator.
[0061] Cycle-consistent loss:
[0062] The cycle consistency loss function with a single generator is defined as:
[0063]
[0064] Loss of identity:
[0065] The identity loss with a single generator is calculated as:
[0066]
[0067] Feature matching loss:
[0068] The feature matching loss with a single generator is calculated as:
[0069]
[0070] in, is the discriminator D x The i-th layer feature extractor of Return from D x The features extracted from the i-th layer, N x is the discriminator D x The total number of layers, Yes D x The number of elements in the i-th level; is the discriminator D y The j-th layer feature extractor of Return from D y The features extracted from the jth layer, N y is the total number of feature layers, Yes D y The number of elements in the jth layer;
[0071] Cycle semantic consistency loss:
[0072] The recurrent semantic label consistency objective with a single generator is calculated as:
[0073]
[0074] in and They are the calibrated true semantic label masks, Φ seg (·) is the semantic segmentation mask generated by the pre-trained semantic segmentation model for the generated image, is the cross entropy loss between the two label masks;
[0075]
[0076] Where N p is the number of pixels, represents the true label of the i-th pixel, Represents the model's predicted value for the i-th pixel.
[0077] A method for detecting a navigable area on a water surface for an unmanned ship, using the above-mentioned water surface reflection and reflection removal method, comprises the following steps:
[0078] S1, obtaining the data of the water surface image in front of the unmanned ship;
[0079] S2. De-noising the water surface image. De-noising is to remove or weaken the water surface reflection and highlight reflection. The removal or weakening is to adopt an image style conversion method to convert the water surface image with a strong water surface reflection style into a water surface image with no reflection or weak reflection style, and convert the water surface image with a strong highlight reflection style into a water surface image with no reflection or weak reflection style;
[0080] S3, inputting the denoised image into a semantic segmentation network, segmenting the water surface, obtaining a water surface classification result, and classifying the water surface image into water surface and non-water surface;
[0081] S4, merging and extracting the image regions classified as water surfaces into water surface connected regions;
[0082] S5. If there is only one water surface connected area, directly use the water surface connected area as the water surface navigable area. If there are multiple water surface connected areas, select the largest connected area that the unmanned ship path can reach as the final unmanned ship water surface navigable area.
[0083] Compared with the prior art, the present invention has the following beneficial effects:
[0084] 1. This paper proposes a two-stage method for segmenting navigable areas on the water surface. In the first stage, the RDGAN network is used to remove or reduce the reflection and interference of the water surface of the river image. In the second stage, the denoised image is transmitted to the semantic segmentation network. The water body segmentation is used to determine the navigable area on the water surface for autonomous cruising of unmanned ships. The problem of semantic pixel misclassification caused by water surface interference is solved;
[0085] 2. By improving the CycleGAN model, the AdaIN layer is added to the generator, and the AdaIN encoding generator (AdaINC) is proposed. AdaINC can automatically generate the normalized affine parameters of the AdaIN layer. AdaINC allows the model to complete training using only a single generator, reducing the network training complexity and memory requirements. In addition, the style of the denoised image can be flexibly controlled by adjusting the AdaINC harmonic parameters. AdaINC allows the generated image to absorb more low-noise water surface features;
[0086] 3. The present invention proposes a CSPADE layer, which is a CSPADE with content and spatial information feature fusion, and CSPADE can retain the semantic level space and content features, and improve the adaptability of content and space in high-dimensional space. It can better eliminate the reflection and reflection of the water surface under harsh environmental conditions, and obtain an enhanced image with richer semantic feature information of the clean water surface.
[0087] 4. The present invention modifies the loss function of the original CycleGAN network, making the network training more stable, making the image after removing reflections and inversions more realistic, and keeping the background unchanged from the original image.
[0088] 5. The present invention simultaneously removes or weakens the water surface reflection and the highlight reflection, and can cope with the problem of complex and changeable water surface environment perception. BRIEF DESCRIPTION OF THE DRAWINGS
[0089] Figure 1 It is the main flow chart of the method for detecting the navigable area on the water surface of the present invention;
[0090] Figure 2 Schematic diagram of the water surface navigable area detection framework based on the two-stage method proposed in this invention
[0091] Figure 3 This is the overall structure diagram of the water surface noise reduction network of the present invention, taking the removal of water surface reflection as an example;
[0092] Figure 4 It is a structural diagram of a generator in the water surface noise reduction network of the present invention;
[0093] Figure 5 It is a flow chart of three normalization modules in the water surface noise reduction network generator of the present invention;
[0094] Figure 6 This is a structural diagram of the calculation method for CSPADE normalization in the water surface noise reduction network generator of the present invention. DETAILED DESCRIPTION
[0095] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0096] See also Figure 1-6 In an embodiment of the present invention, a method for removing reflections and images on a water surface is provided, which uses a deep learning neural network algorithm to reduce noise on a water surface image. The noise reduction is to remove or weaken the reflections and highlight reflections on the water surface. The removal or weakening process is to use an image style conversion method to convert a water surface image with a strong water surface reflection style into a water surface image with no reflection or weak reflection style, and to convert a water surface image with a strong highlight reflection style into a water surface image with no reflection or weak reflection style.
[0097] The method of water surface image style transfer uses a generative adversarial network algorithm, specifically an improved CycleGAN network, called RDGAN. The specific improvements of RDGAN are as follows:
[0098] Step 1: Assume that X and Y are high-noise style source domain images and low-noise style target domain images, respectively. By combining the two generators G in the original cycle-consistency generative adversarial network CycleGAN model X→Y and G Y→X Improved to use only a single generator G with a lightweight AdaIN encoder (AdaINC) implementation;
[0099] Step 2: In the encoder, the adaptive instance normalization (AdaIN) module and the spatial adaptive normalization (SPADE) module are used alternately for normalization, and in the decoder, the adaptive instance normalization (AdaIN) module and the content and spatial adaptive normalization module (CSPADE) are used alternately for normalization;
[0100] Step 3, controlling the quality of the generated water surface denoised image by generating AdaIN harmonic interpolation;
[0101] Step 4: By improving the objective loss function of training optimization in the CycleGAN model, especially adding feature matching loss and cycle semantic consistency loss, the stability of the training model is improved while ensuring that the generated image is more realistic and the background remains unchanged from the original image.
[0102] For the water surface denoising network RDGAN, two types of data sets, high noise and low noise, need to be collected for training. First, the water surface images in the training set are divided into high noise (strong reflection, strong reflection) images and low noise (weak reflection, weak reflection) image sets, which are used as two types of image sets with different styles in the original domain and the target domain respectively. Figure 2 As shown in the figure, in the first stage, these data sets are used to train the water surface image denoising network RDGAN, and then in the second stage, the denoised images are used to train the water body semantic segmentation network. Through these two stages, the training of the water surface navigable area detection model is completed.
[0103] Compared with high-noise images, the overall texture of the water surface in low-noise images is similar, and is almost unaffected during water surface segmentation, and the area where the water surface is located can be completely segmented. We call these low-noise image data clean images, whose water surface reflections and reflections are weak, mainly including flowing water surfaces, dark or turbid water surfaces, raining or just raining water surfaces, etc. When the water flows or the wind blows, the water surface will have a ripple texture, which will weaken the reflected image. When the color of the water body itself is too dark, the water surface will also have no reflective ability. Rainy days are accompanied by strong winds and raindrops, which will form a dripping effect on the water surface. In addition, on rainy days, the mud particles and impurities on the river bank mixed with the water will flow into the river, causing the river water to be turbid, which will make the water surface reflection image disappear or weaken.
[0104] AdaIN in the AdaIN module adaptively generates affine parameters based on the input style encoding, and its normalized calculation formula is as follows:
[0105]
[0106] where μ i (y) and σ i (y) is the mean and variance of the target style, each feature map x i The affine parameters corresponding to the style y are used for scaling and offset to achieve normalization.
[0107] AdaIN aligns the normalized channel mean and variance of the image containing high noise content of the water surface with the mean and variance of the weak noise pattern image so that the generated picture has the same feature distribution as the low noise water area.
[0108] The AdaIN affine parameter K of the AdaIN module is obtained from the AdaIN encoder and is defined as follows:
[0109]
[0110] Where F is the AdaIN code generator, v is the input random Gaussian vector, and F m The mapping network consists of MLP and fully connected layers, X and Y are high-noise style source domain images and low-noise style target domain images respectively. i (y) and σ i (y) is the mean and variance of the low-noise water surface image in the target domain, and is the normalized affine transformation parameter of AdaIN.
[0111] During the forward conversion in the training phase, the AdaIN encoder that converts the high-noise water surface image X into the low-noise water surface image Y is set to a fixed constant vector, specifically composed of a vector with a mean of 0 and a variance of 1, F(v) = K 0 =(0,1)
[0112] In the reverse conversion, the AdaIN encoding of the low-noise water surface image Y into the high-noise water surface image X is generated by the mapping network F m The generated vector is obtained.
[0113] The quality of the generated image is controlled by interpolating the generated AdaIN code. The AdaIN code interpolation calculation method is:
[0114] I(θ)=(1-θ)K 0 +θF m (v),0≤θ≤1
[0115] where I(θ) is between K 0 and F m (v) is the AdaIN harmonic interpolation, and θ is the AdaIN harmonic interpolation coefficient. Larger values of θ indicate more transition to high noise features, while smaller values preserve more low noise features.
[0116] The two generators are improved to use only a single generator G, which is implemented by the AdaIN encoder. The reversible transformation generator from domain X to domain Y is defined as
[0117] G X→Y =G(x,I(0))
[0118] G Y→X =G(y,I(1))
[0119] like Figure 3As shown, only one generator G(x) and AdaIN harmonic interpolation I(.) are used to realize the conversion between the two domains. I(.) is the AdaIN harmonic interpolation.
[0120] CSPADE uses the semantic mask and content code of the image as conditional input to solve the conversion problem from the source domain to the target domain. Affine transformation parameters are obtained by adaptively fusing content information and semantic structure features. The activation value is adjusted by coordinating the scale parameters of the two to retain useful content feature information and better improve the adaptability of content and space. The improved SPADE module, because it also fuses content feature information as input in addition to semantic mask information, is called the content and space adaptive normalization CSPADE module. Figure 6 As shown, the affine parameters and It is the feature map at the i-th layer (n∈N,c∈C i ,y∈H i ,x∈W i ) is combined with the input semantic segmentation mask m x and content encoding c i The specific definition of affine transformation normalization is:
[0121]
[0122] in, is the activation value of the i-th layer, the scaled value and deviation value They are the semantic masks m x and content information c i The weighted sum of is calculated as:
[0123]
[0124] Among them, m x is the semantic mask information of the corresponding layer after resizing, c i It is the content feature encoding of the corresponding scale in the low-dimensional space. and represents the mean and standard deviation and are defined as follows:
[0125]
[0126]
[0127] Where N is the number of samples, H i and W i are the height and width of layer i respectively. In particular, when the weight ω γ = 0, CSPADE will degenerate into the original SPADE, so it is set to ω in the encoder network of the generator during training.γ =0, and in the high-dimensional space of the decoding network, set ω γ >0, obtain the adaptability of content and space in high-level space, and this parameter value is obtained through training.
[0128] Specifically, in the training phase, the semantic mask is provided by the semantic labels annotated in the training set images, while in the inference phase, the semantic mask is given an initial input as an estimate by the pre-trained lightweight FCN network. The semantic information is then corrected and calibrated by the RDGAN network.
[0129] The generator network architecture is as follows Figure 4 As shown, including content encoding module G E , content decoding module G D , and the AdaIN code generation module AdaINC. The encoding and decoding networks fuse low-level and high-level features through jump connections, which can retain more detailed features and improve convergence speed and performance. In the implementation example, at each stage of the encoder part, the width and height of the feature map are halved and the number of channels is doubled. In the encoder part, a stride convolution with a step size of 2 is used to replace the pooling layer. The feature map of the encoder part is jump-connected to the feature map of the decoder part of the corresponding scale at the same stage. In the decoder part, a 2×2 upsampling layer is used. Use an AdaIN layer with a fixed vector or use the output vector from the AdaIN code generator to normalize the feature map
[0130] AdaIN residual block, such as Figure 5 As shown in (a), it consists of convolution, AdaIN, LeakyRelu and skip connection. The AdaIN residual module consists of the AdaIN encoded feature f i As an additional input, f i The feature map is normalized by adjusting the output mean and variance of the activation value as the affine transformation parameters of the AdaIN layer.
[0131] SPADE residual block SPADEResBlock Figure 5 As shown in (b), it consists of convolution, AdaIN, LeakyRelu and skip connection. The additional input of the modulation parameters is the semantic mask m of different scales. x .
[0132] The CSPADE residual block is Figure 5 As shown in (c), it consists of convolution, AdaIN, LeakyRelu and skip connection. The additional input of the modulation parameters includes not only the semantic masks m of different scales x , and also includes content feature encoding at different scales c i .
[0133] AdaINC is a mapping network composed of multiple multi-layer perceptrons (MLPs) and fully connected layers, with input being a randomly sampled 10-bit Gaussian noise vector v of size 1*64. The mapping network learns the noise-free or weak noise style features of the target domain from the vector v through MLP. It consists of multiple fully connected layers, the first few layers are shared parameters, and in order to make the length of the output vector equal to the number of channels of the feature map to be normalized, the last layer is not shared. In the last layer, linear activation is used for the generation of the affine parameter mean vector, and ReLU activation function is used for the generation of the variance to ensure non-negativity.
[0134] The loss function of the improved CycleGAN network training objective is defined as:
[0135]
[0136] in, is a confrontational loss, is the cycle consistency loss, It is a loss of identity. is the feature matching loss, is the cycle semantic consistency loss, λ cyc ,λ ident ,λ fm ,λ sem Represent the weight coefficients of the corresponding losses. Training is done by solving the following minimum-maximum objective function:
[0137]
[0138] (1) Adversarial Loss
[0139] Given a set of unpaired data from a noisy domain X and a noise-free domain Y, by computing the discriminative loss of the generated images and the discriminative loss of the real images after domains X→Y and Y→X, the adversarial loss with a single generator is expressed as follows:
[0140]
[0141] in‖·‖ 2 is the L2 norm, x∈X, x∈Y, G(x) represents the generated low-noise image, D X and D Y is the discriminator.
[0142] (2) Cycle consistency loss
[0143] In order to prevent the image generated by the reverse generator from deviating too much from the target domain image when training the unsupervised water surface image denoising model. For each input image x, the reconstructed image can bring x back to the original image.
[0144]
[0145] Similarly, for each image y from domain Y, the generator G should satisfy that the reverse reconstructed image is also consistent:
[0146]
[0147] The cycle consistency loss function with a single generator is defined as:
[0148]
[0149] (3) Loss of identity
[0150] To prevent the target domain image from being overly distorted due to the transition of the adversarial loss, the identity loss limits this extreme situation by imposing constraints on the algorithm. The identity loss calculates the difference between the input and output. The identity loss calculation formula with a single generator is:
[0151]
[0152] (4) Feature matching loss
[0153] The feature matching loss is used to minimize the difference in high-frequency structure between the original image and the reconstructed image. The high-frequency structure is a feature extracted from the multi-layer structure of the discriminator Dx and Dy. The feature matching loss can ensure that the generated image results are more realistic. The feature matching loss calculation formula with a single generator is:
[0154]
[0155] in, is the discriminator D x The i-th layer feature extractor of Return from D x The features extracted from the i-th layer, N x is the discriminator D x The total number of layers, Yes D x The number of elements in the i-th level; is the discriminator D y The j-th layer feature extractor of Return from D y The features extracted from the jth layer, N y is the total number of feature layers, Yes D y The number of elements in the jth level.
[0156] (5) Cycle semantic consistency loss
[0157] The original image and the denoised image should have semantic label consistency so that the water surface segmentation result of the denoised image remains unchanged in space. The calculation formula for the cyclic semantic label consistency objective with a single generator is:
[0158]
[0159] in and They are the calibrated true semantic label masks, Φ seg (·) is the semantic segmentation mask generated by the pre-trained semantic segmentation model for the generated image, is the cross entropy loss between the two label masks.
[0160]
[0161] Where N p is the number of pixels, represents the true label of the i-th pixel, Represents the model's predicted value for the i-th pixel.
[0162] The discriminator in the implementation example uses a similar structure to the discriminator in PatchGAN. The stride of the first four convolutional layers of the discriminator is 2, and the stride of the remaining convolutional layers is 1. The first convolutional layer takes an input image of one channel and generates a 64-channel feature map. After that, the number of channels doubles each time the feature map passes through the convolutional layer. In the last layer, the final output result is obtained by reducing the number of channels to one, that is, the true or false judgment result of the generated image.
[0163] Implementation Example For the training of RDGAN, the Adam optimizer is used for optimization, and the learning rate is adjusted using the Step strategy. The initial learning rate is set to 0.005, the weight decay is set to 0.0001, the momentum is set to 0.5, the number of iterations is 1000, the batch size is 1, and the crop size is 512×512.
[0164] like Figure 1 As shown, a method for detecting a navigable area on a water surface for an unmanned ship comprises the following steps:
[0165] S1, obtaining the data of the water surface image in front of the unmanned ship;
[0166] S2. De-noising the water surface image. De-noising is to remove or weaken the water surface reflection and highlight reflection. The removal or weakening is to adopt an image style conversion method to convert the water surface image with a strong water surface reflection style into a water surface image with no reflection or weak reflection style, and convert the water surface image with a strong highlight reflection style into a water surface image with no reflection or weak reflection style;
[0167] S3, inputting the denoised image into a semantic segmentation network, segmenting the water surface, obtaining a water surface classification result, and classifying the water surface image into water surface and non-water surface;
[0168] S4, merging and extracting the image regions classified as water surfaces into water surface connected regions;
[0169] S5. If there is only one water surface connected area, directly use the water surface connected area as the water surface navigable area. If there are multiple water surface connected areas, select the largest connected area that the unmanned ship path can reach as the final unmanned ship water surface navigable area.
[0170] Implementation Example For the training of the water surface semantic segmentation model, stochastic gradient descent (SGD) is used for optimization, and the Poly strategy is used to adjust the learning rate. The loss function is set to the cross entropy loss (CE) function, the initial learning rate is set to 0.0001, the weight decay is set to 0.0001, the momentum is set to 0.9, the number of iterations is set to 5000, the batch size is set to 8, and the crop size is set to 512×512.
[0171] It will be apparent to those skilled in the art that the invention is not limited to the details of the exemplary embodiments described above and that the invention can be implemented in other specific forms without departing from the spirit or essential features of the invention. Therefore, the embodiments should be considered exemplary and non-limiting in all respects, and the scope of the invention is defined by the appended claims rather than the foregoing description, and it is intended that all variations falling within the meaning and scope of the equivalent elements of the claims be included in the invention. Any reference numeral in a claim should not be considered as limiting the claim to which it relates.
[0172] In addition, it should be understood that although the present specification is described according to implementation modes, not every implementation mode contains only one independent technical solution. This description of the specification is only for the sake of clarity. Those skilled in the art should regard the specification as a whole. The technical solutions in each embodiment may also be appropriately combined to form other implementation modes that can be understood by those skilled in the art.
Claims
1. A method for removing water surface reflections and reflections, characterized in that: The water surface reflection and highlight reflection are removed or weakened by image style conversion method; The image style transfer method adopts the RDGAN network, and the construction of the RDGAN network includes the following steps: Step 1: X and Y are high-noise style source domain images and low-noise style target domain images, respectively. The two generators G in the original cycle consistency generative adversarial network CycleGAN model are X→Y and G Y→X The improvement is to use only a single generator G and attach a lightweight AdaIN encoder implementation, and the AdaIN encoder uses the AdaINC code generator; Step 2: In the AdaIN encoder, the adaptive instance normalization AdaIN module and the spatial adaptive normalization SPADE module are alternately used for normalization; In the decoder, the adaptive instance normalization AdaIN module and the content and space adaptive normalization CSPADE module are used alternately for normalization. The content and space adaptive normalization module CSPADE uses the semantic mask and content code of the image as conditional input to solve the conversion problem from the source domain to the target domain; the affine transformation parameters are obtained by adaptively fusing the content information and the semantic structure features; the activation value is adjusted by coordinating the scale parameters of the two to retain useful content feature information and improve the adaptability of content and space; Step 3, controlling the quality of the generated water surface denoised image by generating AdaIN harmonic interpolation; Step 4: Modify the objective loss function of training optimization in the CycleGAN model and add feature matching loss and cycle semantic consistency loss.
2. A method for removing water surface reflections and reflections according to claim 1, characterized in that: The AdaIN in the AdaIN module adaptively generates affine parameters according to the input style encoding, and the normalized calculation formula is as follows: where μ i (y) and σ i (y) is the mean and variance of the target style, each feature map x i Use the affine parameters corresponding to style y to perform scaling and offset to achieve normalization; AdaIN aligns the normalized channel mean and variance of images containing high noise content of water surfaces with the mean and variance of weak noise pattern images so that the generated images have the same feature distribution as low-noise water areas.
3. A method for removing water surface reflections and reflections according to claim 2, characterized in that: The AdaIN affine parameter K of the AdaIN module is obtained by the AdaIN encoder and is defined as follows: Where F is the AdaIN code generator, v is the input random Gaussian vector, and F m The mapping network consists of MLP and fully connected layers, X and Y are high-noise style source domain images and low-noise style target domain images respectively; μ i (y) and σ i (y) is the mean and variance of the low-noise water surface image in the target domain, and is the normalized affine transformation parameter of AdaIN; During the forward conversion in the training phase, the AdaIN encoder for converting the high-noise water surface image X to the low-noise water surface image Y is set to a fixed constant vector, specifically composed of a vector with a mean of 0 and a variance of 1, F(v) = K0 = (0, 1); In the reverse conversion, the AdaIN encoding of the low-noise water surface image Y into the high-noise water surface image X is generated by the mapping network F m The generated vector is obtained.
4. A method for removing water surface reflections and reflections according to claim 1, characterized in that: The quality of the generated image is controlled by interpolating the generated AdaIN code. The AdaIN code interpolation calculation method is: I(θ)=(1-θ)K0+θF m (v),0≤θ≤1 where I(θ) is the value between K0 and F m (v) AdaIN harmonic interpolation, θ is the AdaIN harmonic interpolation coefficient; Larger values of θ indicate more transition to high-noise features, while smaller values preserve more of low-noise features.
5. The method for removing water surface reflections and reflections according to claim 1, characterized in that: The two generators in step 1 are improved to use only a single generator G, which is implemented by AdaIN encoder. The reversible transformation generator from domain X to domain Y is defined as G X→Y =G(x,I(0)) G Y→X =G(y,I(1)) Only one generator G(x) and AdaIN harmonic interpolation I(.) are used to achieve the conversion between the two domains; I(.) is the AdaIN harmonic interpolation.
6. A method for removing water surface reflections and reflections according to claim 2, characterized in that: The CSPADE module in step 1 is obtained by fusing content feature information in the SPADE module as input; Affine parameters and It is the joint input semantic segmentation mask m at the i-th layer feature map x and content encoding c i Obtained, where n∈N,c∈C i ,y∈H i ,x∈W i ; The specific definition of affine transformation normalization is: in, is the activation value of the i-th layer, the scaled value and deviation value They are the semantic masks m x and content information c i The weighted sum of is calculated as: Among them, m x is the semantic mask information of the corresponding layer after resizing, c i It is the content feature encoding of the corresponding scale in the low-dimensional space. and represents the mean and standard deviation and are defined as follows: Where N is the number of samples, H i and W i are the height and width of layer i respectively; when the weight ω γ = 0, CSPADE will degenerate into the original SPADE, so it is set to ω in the encoder network of the generator during training. γ =0, and in the high-dimensional space of the decoding network, set ω γ >0, obtain the adaptability of content and space in high-level space, and the parameter value is obtained through training; In the training stage, the semantic mask is provided by the semantic labels annotated in the training set images, while in the inference stage, the semantic mask is given an initial input as an estimate by the pre-trained lightweight FCN network; then the semantic information is corrected and calibrated through the RDGAN network.
7. A method for removing water surface reflections and reflections according to claim 1, characterized in that: The residual block of AdaIN includes convolution, AdaIN, LeakyRelu and skip connection; The AdaIN residual module consists of the AdaIN encoded features f i As an additional input, f i The feature map is normalized by adjusting the output mean and variance after the activation value as the affine transformation parameters of the AdaIN layer; The residual block SPADEResBlock of SPADE includes convolution, AdaIN, LeakyRelu and skip connection, and the additional input of the modulation parameter is the semantic mask m of different scales x ; The residual block of the CSPADE includes convolution, AdaIN, LeakyRelu and skip connection, and the additional input of the modulation parameters includes semantic masks m of different scales. x and content feature encoding at different scales c i .
8. The method for removing water surface reflections and reflections according to claim 1, characterized in that: The AdaINC is a mapping network consisting of multiple multi-layer perceptrons (MLPs) and fully connected layers. The input of the mapping network is a randomly sampled 10-bit Gaussian noise vector v of size 1*64. The mapping network learns the noise-free or weak noise style features of the target domain from the vector v through the MLP. The mapping network consists of multiple fully connected layers, the last layer of the mapping network is not shared, and the remaining layers share parameters. In the last layer, linear activation is used to generate the affine parameter mean vector, and ReLU activation function is used to generate the variance to ensure non-negativity.
9. The method for removing water surface reflections and reflections according to claim 1, characterized in that: Improve the stability of the training model by improving the loss function of the CycleGAN model, which includes feature matching loss and cycle semantic consistency loss The loss function of the CycleGAN network training objective is defined as: in, is a confrontational loss, is the cycle consistency loss, It is a loss of identity. is the feature matching loss, is the cycle semantic consistency loss, λ cyc ,λ ident ,λ fm ,λ sem Respectively represent the weight coefficients of the corresponding losses; Training is done by solving the following min-max objective function: Adversarial Loss: The adversarial loss with a single generator is expressed as follows: where ‖·‖2 is the L2 norm, x∈X, x∈Y, G(x) represents the generated low-noise image, and D X and D Y is the discriminator; Cycle-consistent loss: The cycle consistency loss function with a single generator is defined as: Loss of identity: The identity loss with a single generator is calculated as: Feature matching loss: The feature matching loss with a single generator is calculated as: in, is the discriminator D x The i-th layer feature extractor of Return from D x The features extracted from the i-th layer, N x is the discriminator D x The total number of layers, Yes D x The number of elements in the i-th level; is the discriminator D y The j-th layer feature extractor of Return from D y The features extracted from the jth layer, N y is the total number of feature layers, Yes D y The number of elements in the jth layer; Cycle semantic consistency loss: The recurrent semantic label consistency objective with a single generator is calculated as: in and They are the calibrated true semantic label masks, Φ seg (·) is the semantic segmentation mask generated by the pre-trained semantic segmentation model for the generated image, is the cross entropy loss between the two label masks; Where N p is the number of pixels, represents the true label of the i-th pixel, Represents the model's predicted value for the i-th pixel.
10. A method for detecting navigable areas on the water surface for unmanned ships, using the water surface reflection and reflection removal method according to any one of claims 1 to 9, characterized in that: The following steps are involved: S1, obtaining the data of the water surface image in front of the unmanned ship; S2. De-noising the water surface image. De-noising is to remove or weaken the water surface reflection and highlight reflection. The removal or weakening is to adopt an image style conversion method to convert the water surface image with a strong water surface reflection style into a water surface image with no reflection or weak reflection style, and convert the water surface image with a strong highlight reflection style into a water surface image with no reflection or weak reflection style; S3, inputting the denoised image into a semantic segmentation network, segmenting the water surface, obtaining a water surface classification result, and classifying the water surface image into water surface and non-water surface; S4, merging and extracting the image regions classified as water surfaces into water surface connected regions; S5. If there is only one water surface connected area, directly use the water surface connected area as the water surface navigable area. If there are multiple water surface connected areas, select the largest connected area that the unmanned ship path can reach as the final unmanned ship water surface navigable area.