Pseudo-sonar underwater dam body crack image generation method based on style conversion
Through the style-transformed pseudosonar underwater dam body crack image generation method, the marked ground crack image data set generates high-quality underwater sonar dam body crack pseudo-image, which solves the problem of poor image recognition effect of underwater sonar dam body crack image, and achieves more accurate dam crack detection and recognition.
Patent Information
- Application Number
- CN202510152178.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-12
- Publication Date
- 2025-05-23
AI Technical Summary
In the prior art, the image recognition effect of underwater sonar dam body cracks is poor, mainly due to poor sonar image quality and insufficient data set of marked underwater sonar dam body cracks, resulting in overfitting of model training.
The style-conversion-based pseudosonar underwater dam body crack image generation method is adopted, and the marked ground crack image data set is used to generate realistic underwater dam body crack pseudo-images, through the specific domain characteristics of the content image, the style-conversion pseudo-image generation network based on image patches, and the fast-guided filter enhancement module to generate high-quality pseudo-sonar images.
The demand for labeling underwater sonar dam fracture samples is significantly reduced, the training effect of the identification model is improved, more accurate dam fracture detection is achieved, the labor demand for data labeling is reduced, and the accurate identification of dam fractures is improved.
Smart Images

Figure CN120032004A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of image generation by style conversion, and in particular to a method for generating a pseudo-sonar underwater dam crack image based on style conversion. Background Art
[0002] As a key component of water conservancy project construction, dams have far-reaching significance in agricultural irrigation, hydropower generation and disaster prevention; however, as the operation time of reservoir dams continues to increase, they are usually accompanied by a series of problems such as changes in the geological environment, degradation of drainage and anti-seepage functions, all of which may cause cracks in the dam body. Traditional artificial methods for detecting dam body cracks, such as electrical exploration, elastic wave testing, tomography and ground penetrating radar, are expensive, inconvenient and unreliable. Sonar image detection of underwater dam body cracks is non-destructive, intuitive, convenient and efficient. However, due to the limitations of sonar acquisition equipment and the influence of the underwater environment of the dam, underwater sonar dam body crack images contain very little effective information. At present, most underwater dam body crack recognition methods are based on labeled open source ground crack image datasets, and there are very few or no labeled underwater sonar dam body crack image datasets. However, poor sonar image quality and insufficient datasets will cause overfitting of model training, resulting in poor recognition effect of underwater sonar dam body crack images, making it difficult to accurately identify dam body cracks in actual underwater dam body detection. Summary of the invention
[0003] In view of the problems in the prior art, the present invention provides a method for generating pseudo sonar underwater dam crack images based on style transfer, which aims to generate a large number of realistic underwater sonar dam crack pseudo images by using a labeled ground crack image dataset, significantly reducing the demand for labeled underwater dam crack samples, improving the training effect of the recognition model, achieving more accurate dam crack detection, and reducing the labor required for data annotation and the demand for sonar dam crack data, while improving the accurate identification of dam cracks.
[0004] A method for generating a pseudo-sonar underwater dam crack image based on style transfer comprises the following steps:
[0005] Step 1: Construct an underwater dam crack pseudo image generation network, which includes a content image specific domain feature weakening module, an image patch-based style transfer pseudo image generation network, and a fast guided filtering enhancement module;
[0006] Step 2: Send the optical crack image as the content image data set to the underwater dam body crack pseudo image generation network; the underwater dam body crack pseudo image generation network processes the content image data set as follows:
[0007] Step 2.1: Use the content image specific domain feature weakening module to weaken the non-specific domain features of the content image dataset to obtain the optical content image dataset after feature weakening;
[0008] Step 2.2: Use the side-scan sonar image as a style image dataset, and input the style image dataset and the attenuated optical content image dataset into an image patch-based style transfer pseudo image generation network, and obtain the content image patch sequence of the optical content image dataset and the style image patch sequence of the style image dataset through the image patch-based style transfer pseudo image generation network;
[0009] Step 2.3: Reconstruct the initial pseudo sonar underwater dam crack image using the content image patch sequence and the style image patch sequence;
[0010] Step 2.4: The initial pseudo-sonar underwater dam crack image is used as the guide image and the filtered input image, and the fast guided filter enhancement module is used to perform noise reduction and enhancement;
[0011] Step 3: Generate the final pseudo-sonar underwater dam crack image.
[0012] Further: The process of obtaining the content image patch sequence is:
[0013] Step 2.2.1: Segment the optical crack image in the optical content image dataset into several image patches through PatchEmbed;
[0014] Step 2.2.2: The image block is linearly projected and SCPE module is used to obtain a content image patch sequence with position coding; wherein, firstly, a convolution layer and an adaptive average pooling layer are used to perform an average pooling operation on the content image of the image block; the adaptive average pooling layer pools the input feature map into a fixed-size output; the pooled feature map is linearly projected through a convolution layer, and the convolution layer maps the input feature map from 512 channels to 512 channels to generate a content-related position coding; the generated content position coding is adjusted to the same height and width as the style image through bilinear interpolation; finally, the position coding is added to the content features of the content image as the position information of the content features, and a content image patch sequence combined with the position coding is obtained;
[0015] At the same time, the process of obtaining the style image patch sequence is:
[0016] Step 2.2.3: Use PatchEmbed to split the side-scan sonar image in the style image dataset into several image patches;
[0017] Step 2.24: The image block is linearly projected to obtain a style image patch sequence.
[0018] Further: image block Semantic content-adaptive positional encoding for:
[0019] in, is the average pooling function, is a learnable position encoding function Convolution operation, is a sequence of content image patches A learnable positional encoding of , is the interpolation weight and N is the number of neighboring blocks.
[0020] Further: Step 2.3 is specifically:
[0021] Step 2.3.1: The content image patch sequence and the style image patch sequence are fed to the content image encoder and the style image encoder respectively;
[0022] Step 2.3.2: The outputs of the content image encoder and the style image encoder are fed into a multi-layer decoder;
[0023] Step 2.3.3: The multi-layer decoder translates the encoded content image patch sequence in a regression manner according to the encoded style image patch sequence, and obtains the initial pseudo sonar underwater dam crack image.
[0024] Further: The processing process of the multi-layer decoder is:
[0025] The input to the multi-layer decoder consists of a sequence of encoded content image patches and a sequence of style image patches , generate queries (Q) using a sequence of content image patches, and generate keys (K) and values (V) using a sequence of style image patches:
[0026] The calculation process of the output sequence G of the multi-layer decoder is:
[0027] Among them, the output sequence G of the multi-layer decoder is The shape of the output sequence is not directly upsampled to construct the final result, but a 3-layer CNN decoder is used to refine the subsequent output of the transformer decoder; for each layer, a series of operations are used to expand the scale, including ; This operation converts the output of the multi-layer decoder into a format suitable for convolution operations; the first layer of convolution equals the number of input channels to the embedding dimension and halves the number of output channels, gradually reducing the complexity of the feature space, and using The convolution kernel captures local area features and keeps the output resolution unchanged; the second convolution layer receives the output of the previous layer, further reducing the number of channels and maintaining the same resolution; the last layer reconstructs the resolution Initial pseudo-sonar underwater dam crack image.
[0028] Further: the content image specific domain feature weakening module includes an encoder, a decoder, random noise, a feature whitening module, a Haar wavelet decomposition module and an image reconstruction module, the encoder is VGGEncoder; the decoder is VGGDecoder; the feature whitening module performs whitening transformation in the feature space and redistributes the features through principal component analysis; the Haar wavelet decomposition module decomposes the features into different frequency components for analysis or reconstruction, and obtains low-frequency components and high-frequency components by performing Haar wavelet decomposition, which are passed to the decoder to generate low-frequency feature maps and high-frequency feature maps respectively; the image reconstruction module reconstructs the input image after whitening transformation and Haar wavelet decomposition to generate an image for the features; at the same time, the final results of the whitening operation, noise processing and feature transformation are combined with pixel reconstruction loss and feature loss to achieve reconstruction and generation of feature weakened images.
[0029] Further: Step 2.1 is described by the formula:
[0030] in, Represents the input image The predicted output of It is a feature extraction encoder that extracts deep features from the input image; It is the function responsible for feature separation, which removes the texture and color features in the image and retains the domain-invariant features; It is a classifier that realizes the mapping between deep features and predictions; A set of instances representing content images; A collection of instances representing sonar crack images;
[0031] Then use the pre-trained autoencoder to extract the input image features, and reconstruct the input image using pixel reconstruction loss and feature loss;
[0032] in, The total loss for feature-weakened image reconstruction; It is a measure of the difference between the reconstructed image and the input image, which is the pixel reconstruction loss; It is a measure of the difference between the deep features of the reconstructed image and the input image, which is the feature loss; represents the reconstructed output, represents the input image; It is an encoder that extracts deep features of the image;
[0033] Then perform whitening transformation:
[0034] in, The characteristic value is The diagonal matrix of the covariance matrix, is the orthogonal matrix corresponding to the eigenvectors, represents the whitened content image features, is the content image depth feature;
[0035] Then, the style features extracted by the encoder from the image are fused with the whitened content features through coloring transformation, expressed as formula (4):
[0036] in, Indicates the dyeing characteristics. The characteristic value is The diagonal matrix of the covariance matrix, is the orthogonal matrix corresponding to the eigenvectors, is the style image depth feature, and finally the merged depth feature Decoded by the decoder to reconstruct the image .
[0037] Further: Step 2.4 is described by the formula:
[0038] in, represents the index of the pixel in the guided filter; The output image is pixel value; is the first pixel value; represents a local window with radius r, Display Window The index of , Display Window The linear coefficients in; given an image Under the input condition, the above formula (22) can be obtained:
[0039] in, , To guide the image In the window The variance and mean of represents the regularization parameter related to the degree of smoothness;
[0040] Since each window is calculated , , but each pixel is contained in multiple windows, and for each pixel, multiple , , multiple , The calculated results are averaged to get the output value. The above process is described as follows:
[0041] in, , represents the average coefficient of all windows, and the window is As the center; reduce the number of pixels by downsampling and calculate , The average value is then upsampled to restore to the original size; in high variance areas, The value of is large, the output image It depends on the guide image , thereby retaining edge information; while in low variance areas, The value of is small, the output image Then smoothing is performed; the final pseudo-sonar underwater dam crack image is the output image .
[0042] The beneficial effects of the present invention are as follows: first, the pseudo sonar image generated by the content image specific domain feature weakening module is closer to the real sonar image; then, a style conversion pseudo image generation network based on image patches is used, and the content sequence is stylized according to the style sequence, thereby generating a stylized result of the input content image with good structure and details; finally, filtering or edge preservation processing is performed under different variances by the fast guided filtering enhancement module, thereby generating a pseudo image that is more similar to the real underwater sonar dam crack image; thereby achieving more accurate dam crack detection, reducing the labor required for data annotation and the demand for sonar dam crack data, and improving the accurate identification of dam cracks. BRIEF DESCRIPTION OF THE DRAWINGS
[0043] Figure 1 It is a general framework diagram of constructing a pseudo image generation network for underwater dam cracks in the present invention; Figure 2 The structural diagram of the module for weakening domain-specific features of content images; Figure 3 This is the result diagram of optical crack feature weakening; Figure 4 Generate network graphs for style transfer pseudo images based on image patches; Figure 5 This is the structure diagram of the transformer decoder layer; Figure 6 Schematic diagram of the semantic content adaptive position encoding module; Figure 7 This is the result of generating the pseudo-sonar underwater dam crack image. DETAILED DESCRIPTION
[0044] The present invention is described in detail below in conjunction with the accompanying drawings. The embodiments of the present invention are described in detail below, and examples of the embodiments are shown in the accompanying drawings, wherein the same or similar reference numerals throughout represent the same or similar elements or elements with the same or similar functions. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain the present invention, and cannot be interpreted as limiting the present invention. The directional terms such as left, middle, right, top, and bottom in the embodiments of the present invention are only relative concepts or are based on the normal use state of the product, and should not be considered as restrictive.
[0045] A method for generating a pseudo-sonar underwater dam crack image based on style transfer comprises the following steps:
[0046] Step 1: Construct a pseudo image generation network for underwater dam cracks, such as Figure 1As shown in the figure, the underwater dam crack pseudo image generation network includes a content image specific domain feature weakening module (SDFWMCI), an image patch-based style transfer pseudo image generation network (IPSTPG) and a fast guided filter enhancement module (FGFE);
[0047] Step 2: Send the open-source optical crack images as content image datasets to the underwater dam crack pseudo image generation network; the underwater dam crack pseudo image generation network processes the content image dataset as follows:
[0048] Step 2.1: Use the content image specific domain feature weakening module to weaken the non-specific domain features of the content image dataset to obtain the optical content image dataset after feature weakening;
[0049] Among them, side-scan sonar images have great differences in pixel distribution and image texture compared with traditional optical images, that is, there is a large domain gap between the two images; research has found that the target is deformed after being imaged by the sonar instrument, but compared with the corresponding optical image, the contour features of the side-scan sonar imaging object remain basically unchanged; the present invention selects this unchanging contour feature as the invariant domain feature, and uses texture features such as color and pixels as specific domain features; while keeping the contour feature unchanged, the optical crack image is used as the content source domain, and the real side-scan sonar image is used as the target domain, so as to narrow the domain gap between the two, and combine Figure 2 As shown in the figure, the content image specific domain feature weakening module includes an encoder, a decoder, random noise, a feature whitening module, a Haar wavelet decomposition module and an image reconstruction module. The encoder is VGGEncoder, which includes a convolution layer: performs feature extraction and gradually increases the number of channels to capture more complex features; a filling layer: uses reflective edge filling to reduce the impact of edge effects on the results; an activation function: introduces nonlinear capabilities to the network; a pooling layer: reduces the spatial dimension of the feature map by downsampling and returns the maximum value index for use by the decoder;
[0050] The decoder is the VGGDecoder of the VGG model, which includes convolutional layers: gradually reducing the number of channels to reconstruct the input feature map; non-pooling layers: reversely recovering the spatial resolution of the feature map according to the pooling index returned by the encoder; padding layers and activation functions: helping to restore the spatial distribution while retaining feature information; reversely restoring the encoder feature map to generate a weakened image, and can also combine random noise to simulate the underwater environment;
[0051] The feature whitening module performs whitening transformation in the feature space and redistributes features through principal component analysis. Its main working steps are: centering the features (removing the mean); calculating the covariance matrix and solving the singular value decomposition; applying the whitening matrix to adjust the feature distribution to make it more uniform; after the whitening module, the correlation between features can be reduced, providing more robust data for subsequent generation and analysis;
[0052] The Haar wavelet decomposition module decomposes the features into different frequency components for analysis or reconstruction. By performing Haar wavelet decomposition, low-frequency components and high-frequency components are obtained and passed to the decoder to generate low-frequency feature maps and high-frequency feature maps respectively;
[0053] The image reconstruction module reconstructs the input image after whitening transformation and Haar wavelet decomposition to generate an image with features. At the same time, the final results of whitening operation, noise processing and feature transformation are combined to achieve the reconstruction generation of feature-weakened image by using pixel reconstruction loss and feature loss.
[0054] Step 2.1 is described by the formula:
[0055] in, Represents the input image The predicted output of It is a feature extraction encoder that extracts deep features from the input image; is a function responsible for feature separation, which is used to remove features such as texture and color in the image and retain domain-invariant features (such as contour features); in the present invention, This is achieved through whitening transformation, which aims to reduce the domain difference between the source domain and the target domain; It is a classifier that realizes the mapping between deep features and predictions, specifically the softmax function; A set of instances representing content images; A collection of instances representing sonar crack images;
[0056] Then use the pre-trained autoencoder to extract the input image features, and reconstruct the input image using pixel reconstruction loss and feature loss;
[0057] in, The total loss for feature-weakened image reconstruction; It is a measure of the difference between the reconstructed image and the input image, that is, the pixel reconstruction loss, which is used to ensure that the network can retain the information of the input image as much as possible; It is a measure of the difference between the deep features of the reconstructed image and the input image, that is, the feature loss, which ensures that the network can maintain consistency in the feature space; represents the reconstructed output, represents the input image, specifically an image in the content image dataset; It is an encoder for extracting deep features of images, specifically VGGEncoder;
[0058] Then perform whitening transformation:
[0059] in, The characteristic value is The diagonal matrix of the covariance matrix, is the orthogonal matrix corresponding to the eigenvectors, represents the whitened content image features, is the content image depth feature;
[0060] Then, the style features extracted by the encoder from the image are fused with the whitened content features through coloring transformation, expressed as formula (4):
[0061] in, Indicates the dyeing characteristics. The characteristic value is The diagonal matrix of the covariance matrix, is the orthogonal matrix corresponding to the eigenvectors, is the style image depth feature, and finally the merged depth feature Decoded by the decoder to reconstruct the image .
[0062] In this task, when only the contour features are retained, the efficiency of style transfer can be improved; if the whitening transform is directly used as a feature weakening network, the reconstructed image still has residual style features and some colors; in order to remove redundant features, this paper decomposes the whitened deep features by Harr wavelet transform, and low-frequency components and high-frequency components can be obtained after decomposition; since the length of the low-frequency component and the high-frequency component is half of the depth feature, this paper uses nearest neighbor interpolation to restore them to the same length as the depth feature; then the decoder is used to reconstruct the two decomposed features; however, in the low-frequency component, the color and texture features become clearer, which can be regarded as the result of low-pass filtering of the whitened deep feature, and the domain gap after filtering becomes larger; and adding noise to the whitened deep feature greatly weakens the decoder of the texture feature of the output result, thereby reducing the difference between the source domain and the target domain; the addition of noise pollution can reduce the domain gap, but the removal effect of the texture feature is still insufficient; this paper alternately uses non-pooling layers and upsampling layers to expand the feature map in the decoder, and the relevant details are as follows:
[0063] In the forward function, the fourth layer of unpooling (unpool3) restores the size of the feature map obtained from the fourth layer of maximum pooling (maxpool3) of the encoder; in the Harr wavelet transform, the feature map is enlarged to twice the spatial size through the upsampling layer to further restore the details; in general, unlike the traditional commonly used upsampling layer, the unpooling layer uses the maximum activation position saved during the pooling operation as an indicator and uses them to put each activation back to its original pooling position; the texture information is retained through the maximum activation value position indicator of the maximum pooling layer; the unpooling layer is the inverse operation of the maximum pooling, which uses the index information saved during pooling and the pooling The spatial distribution of the feature map is restored by resizing the feature map after the feature map is enlarged; the non-pooling and pooling operations are coordinated to restore the spatial information lost in the pooling process; the present invention uses a traditional upsampling layer at an appropriate position. The upsampling layer enlarges the feature map in a way that ignores the active position. Although it may destroy the transmission of feature information, it avoids the disadvantage of using the unpooling layer to retain too much information; the upsampling layer uses bilinear interpolation to enlarge the feature map without relying on a specific feature position or pooling index; the upsampling layer uses interpolation in the decoder to increase the spatial resolution and restore the size of the image; the optical content image specific field feature weakening module structure, such as Figure 2 As shown;
[0064] In order to better achieve feature weakening, this paper introduces random noise during whitening transform and wavelet decomposition. The random noise can be directly used as input for decoder processing. By using pure random noise, the robustness of the decoder to meaningless input features or the sensitivity of the output characteristics can be tested, while improving the effect of feature weakening.
[0065] The details of the real underwater sonar dam crack image are not rich enough, but from the reconstruction results, we know that the color and texture of the low-frequency feature image after decoding and reconstruction are very clear. Too much detail will affect the quality of the generated image; the high-frequency features after decoding have weakened most of the style features and colors, retaining the main structure of the image, and obtaining a better feature weakened content image; the optical crack image feature weakening results, such as Figure 3 As shown;
[0066] Step 2.2: Use the side-scan sonar image as a style image dataset, and input the style image dataset and the attenuated optical content image dataset into an image patch-based style transfer pseudo image generation network, and obtain the content image patch sequence of the optical content image dataset and the style image patch sequence of the style image dataset through the image patch-based style transfer pseudo image generation network;
[0067] Among them, combined Figure 4 As shown in Figure 2, the process of obtaining the content image patch sequence is as follows:
[0068] Step 2.2.1: Segment the optical crack image in the optical content image dataset into several image patches through PatchEmbed;
[0069] In order to eliminate the biased representation problem of the convolutional neural network-based style transfer method, this paper takes advantage of the transformer's strong representation ability to capture accurate content representation and avoid the loss of fine details. A style transfer method based on image patches is designed to generate pseudo images. First, the optical content image and the side-scanning sonar style image with weakened features in the previous stage are divided into small blocks through PatchEmbed, and the size of each image block is ; A two-dimensional convolutional layer is used in PatchEmbed, which is used to divide the image into blocks and embed each small block into a high-dimensional space. Each small block will be mapped into a 768-dimensional vector; then all dimensions of the tensor starting from dimension 2 are flattened. The flattening operation is to arrange vectors of each size into a long vector, so that the features of each image block become a vector, which is convenient for feeding into the subsequent network; the divided image blocks are linearly projected to obtain the patch sequence of content and style image blocks; and the image blocks are mapped to the sequential feature map, and the shape of the feature map is , the length is , where m represents the size of the block, which is set to 8; C represents the dimension;
[0070] When using a transformer-based model, it is necessary to include positional encoding (PE) in the input sequence to obtain structural information; the attention score of the j-th image block and the l-th image block is calculated as:
[0071] in, and is the parameter matrix used for query (Q) and key (K) calculations, represents the j-th one-dimensional PE, represents a sequence of patches; in the two-dimensional case, pixels With pixels The relative position relationship between the patches at the two positions is:
[0072] in, , d=512; the relative position relationship between two image blocks depends only on the spatial distance between them;
[0073] This paper does not directly adopt the classic sine / cosine position encoding, because the traditional position encoding is designed for logically ordered sentences, while image patches are organized according to content. The distance between two patches is expressed as , (red and green patches) and The difference between (the red and blue patches) should be small, e.g. Figure 6 As shown on the right, we expect similar content patches to have similar stylized results; Figure 6 As shown on the left, when the image is scaled, the relative distance between image patches at the same position will change dramatically, which may not be suitable for multi-scale methods in visual tasks. To this end, this paper proposes a semantic content adaptive position encoding (SCPE module), which is scale-invariant and more suitable for style transfer tasks. Unlike the sinusoidal PE that only considers the relative distance of image patches, SCPE is conditioned on the semantics of image content. Assuming that The position encoding of is sufficient to represent the semantics of the image; for optical content images , the fixed The positional encoding is rescaled to ,like Figure 6 As shown on the right; in this way, different image scales will not affect the spatial relationship between the two image blocks; the content image patch sequence with SCPE set is transmitted to the content image encoder, and the style image patch sequence is transmitted to the style image encoder;
[0074] Image Blocks Semantic content-adaptive positional encoding for:
[0075] in, is the average pooling function, is a learnable position encoding function Convolution operation, is a sequence of content image patches A learnable positional encoding of , is the interpolation weight, and N is the number of neighboring blocks. Add to As the jth block at pixel position The final feature embedding at ;
[0076] Step 2.2.2: The image block is linearly projected and SCPE module (semantic content adaptive position encoding) to obtain a content image patch sequence with position encoding; wherein, firstly, a convolution layer and an adaptive average pooling layer are used to perform an average pooling operation on the content image of the image block; the adaptive average pooling layer pools the input feature map into an output of a fixed size (18x18); the pooled feature map is linearly projected through a convolution layer, which maps the input feature map from 512 channels to 512 channels. The purpose of this step is to further extract features from the pooled feature map and generate content-related position encoding; the generated content position encoding is adjusted to the same height and width as the style image through bilinear interpolation; finally, the position encoding is added to the content features of the content image as the position information of the content features, and a content image patch sequence combined with the position encoding is obtained;
[0077] At the same time, the process of obtaining the style image patch sequence is:
[0078] Step 2.2.3: Use PatchEmbed to split the side-scan sonar image in the style image dataset into several image patches;
[0079] Step 2.24: The image block is linearly projected to obtain a style image patch sequence;
[0080] Step 2.3: Reconstruct the initial pseudo-sonar underwater dam crack image using the content image patch sequence and the style image patch sequence; specifically:
[0081] Step 2.3.1: The content image patch sequence and the style image patch sequence are respectively fed to the content image encoder and the style image encoder;
[0082] Step 2.3.2: The outputs of the content image encoder and the style image encoder are fed into a multi-layer decoder;
[0083] Input content patch sequence , first feed it into the content image encoder. Each layer of the content image encoder consists of a multi-head self-attention module (MSA) and a feed-forward network (FFN). The input sequence is encoded into query (Q), key (K), and value (V):
[0084] in, , the multi-head attention calculation formula is as follows:
[0085] in, is a learnable parameter, , Y is the number of attention heads;
[0086] Use residual connection to get the encoded content patch sequence :
[0087] in, , layer normalization is applied after each block;
[0088] Step 2.3.3: The multi-layer decoder translates the encoded content image patch sequence in a regression manner according to the encoded style image patch sequence, and obtains the initial pseudo sonar underwater dam crack image; wherein the processing process of the multi-layer decoder is:
[0089] Input style image patch sequence , and the same operation is used to encode it into a style image patch sequence , but with the content image patch sequence The difference is that the sequence does not apply the position encoding operation, because there is no need to maintain the structure of the input style in the final output; after the above steps, a multi-layer decoder is used to perform sonar image stylization on the content image patch sequence; the multi-layer decoder in this process is used to encode the style image patch sequence according to the encoded Translating encoded content image patch sequences in a regressive manner ; Different from the autoregressive process in NLP tasks, the present invention takes all sequence blocks as input at one time to predict the output; Figure 5 As shown, each transformer decoding layer contains 2 MSA layers and 1 FFN; the input of the multi-layer decoder includes the encoded content image patch sequence and a sequence of style image patches , generate queries (Q) using a sequence of content image patches, and generate keys (K) and values (V) using a sequence of style image patches:
[0090] The calculation process of the output sequence G of the multi-layer decoder is:
[0091] Among them, the output sequence G of the multi-layer decoder is Instead of directly upsampling the output sequence to construct the final result, a 3-layer CNN decoder is used to refine the subsequent output of the multi-layer decoder; for each layer, a series of operations are used to expand the scale, including ; This operation converts the output of the multi-layer decoder into a format suitable for convolution operations; the first layer of convolution equals the number of input channels to the embedding dimension and halves the number of output channels, gradually reducing the complexity of the feature space, and using The convolution kernel captures local area features and keeps the output resolution unchanged; the second convolution layer receives the output of the previous layer, further reducing the number of channels and maintaining the same resolution; the last layer reconstructs the resolution Initial pseudo-sonar underwater dam crack image;
[0092] In order to make the generated pseudo image better have both the structure and sonar style characteristics of the original optical target; the present invention introduces two different perceptual loss terms to measure the initial output image With the input content image The content difference between , and the initial output image With the input style image style differences between; content loss The calculation process:
[0093] in, is the feature extracted at the jth layer, M represents the number of layers;
[0094] Style Loss The calculation process:
[0095] in, represents the mean of the extracted features, The variance of the corresponding extracted features;
[0096] The present invention also adopts identity loss to learn richer and more accurate content and style representations; specifically, two identical content (or style) images are fed into the pseudo image generation network, and the generated output image (or ) should be the same as the input content image (or style image ) are the same; two identity losses are calculated to measure (or )and) (or ) between:
[0097] The entire pseudo image generation network (IPSTPG) is optimized by minimizing the following function to obtain the total loss :
[0098] in, , , , is the weight parameter used to balance various types of losses;
[0099] Step 2.4: The initial pseudo-sonar underwater dam crack image is used as the guide image and the filtered input image, and the fast guided filter enhancement module is used to perform noise reduction and enhancement;
[0100] Since this design adds random noise in the feature weakening stage, it promotes the weakening of the optical crack content image features, and at the same time simulates the underwater noise environment of the dam body by means of the introduced noise; the influence of noise on many generated images is inevitable, and may reduce the quality and visual effect of the generated images; and it is difficult to obtain a clean reference true value image for the generated pseudo image, so it is impossible to use the clean image as a reference to enhance the noisy initial pseudo sonar underwater dam body crack image; this paper proposes a fast guided filtering enhancement module that can balance the noise distribution in the generated image; this design uses the initial pseudo sonar underwater dam body crack image as the guide image and the filter input image for noise reduction and enhancement; the fast guided filtering includes the guide image , filter the input image , a local linear model of the filtered output image ; Among them, the guiding image Provide local feature information; input image is the target image to be processed; sliding window It is used to calculate the local mean and statistical information; regularization is used to balance the filtering effect. The scaling ratio s is used to control the reduction factor of downsampling and upsampling; the workflow is: first, downsample the guide image and the target image to reduce the amount of calculation by reducing the total number of pixels; secondly, calculate the mean, variance, covariance and linear coefficient on the downsampled image; then upsample to restore the linear coefficient to the original image size; finally, calculate the filtering output based on the linear coefficient and the original guide image;
[0101] Step 2.4 is described by the formula:
[0102] in, represents the index of the pixel in the guided filter; The output image is pixel value; is the first pixel value; represents a local window with radius r, Display Window The index of , Display Window The linear coefficients in; given an image Under the input condition, the above formula (22) can be obtained:
[0103] in, , To guide the image In the window The variance and mean of represents the regularization parameter related to the degree of smoothness;
[0104] Since each window is calculated , , but each pixel is contained in multiple windows, and for each pixel, multiple , , multiple , The calculated results are averaged to get the output value. The above process is described as follows:
[0105] in, , represents the average coefficient of all windows, and the window is As the center; reduce the number of pixels by downsampling and calculate , The average value is then upsampled to restore to the original size; in high variance areas, The value of is large, the output image Will rely (more) on the guiding image , thereby retaining edge information; while in low variance areas, The value of is small, the output image Then smoothing is performed;
[0106] In addition, the image is scaled down by downsampling during fast guided filtering, which reduces the computational complexity and is suitable for large-resolution image processing. At the same time, downsampling reduces the total number of pixels, but retains the global feature information, and the output effect can be approximated after upsampling. The range of the scaling factor s is (0, 1], and the image is scaled down to process fewer pixels, which reduces the time complexity. Assume that B represents the total number of pixels in the original image, and the number of pixels after scaling down is , B = image width × image height = W × H; when the fast guided filter is working, when the variance is small and when the guided image is the input image itself, the output image is the image after filtering, that is, the filtering effect is achieved; when the variance is large, the edge preservation effect is achieved;
[0107] Step 3: Generate the final pseudo-sonar underwater dam crack image. The final pseudo-sonar underwater dam crack image is the output image. .
[0108] In order to make the generated pseudo sonar images closer to real sonar images, this paper proposes a module for weakening specific domain features of optical content images, namely the SDFWMCI module. Through a detailed analysis of the characteristics and differences between optical images and sonar images, it is found that although the shape of the target is deformed in the sonar image, the contour features of the target remain basically unchanged as in the conventional optical image, while there are large differences in features such as pixel distribution and texture features. We can select the target contour features of the labeled open source optical ground crack image dataset as domain-invariant features, and its texture features such as color and pixel value distribution information as specific domain features. By maintaining the domain-invariant features, the optical ground crack image can be used as the content source domain, the side-scanning sonar image can be used as the target domain, and random noise can be introduced to narrow the domain gap, which helps to generate more realistic images.
[0109] In order to generate stylized results of input content images with good structure and details, this paper proposes an image patch-based style transfer pseudo image generation network, namely the IPSTPG network. Traditional style transfer methods use convolutional neural networks to learn style and content representations. Due to the limited receptive field of convolution operations, CNNs cannot capture long-distance dependencies when the number of layers is insufficient. However, the increase in network depth may lead to the loss of feature resolution and fine details, which in turn leads to the damage of stylized results in terms of content structure preservation and style presentation. The image patch-based style transfer pseudo image generation network proposed in this paper makes full use of the transformer's ability to capture long-distance dependencies of image features, captures accurate content representation, and avoids the loss of fine details. At the same time, a semantic content-adaptive position encoding (SPCE) is designed and added to the input content image patch sequence to obtain structural information. Traditional position encoding is designed for logically arranged sentences, while image patches are organized according to content. The position encoding designed in this paper expects similar content patches to have similar stylized results. This paper splits the content and style images into patches and uses linear projection to obtain patch sequences. Then, the content sequence with SPCE added is fed into the content transformer encoder, while the style sequence is fed into the style multi-layer encoder. After the two multi-layer encoders, a multi-layer transformer decoder is used to stylize the content sequence according to the style sequence. Finally, a progressive upsampling decoder is used to obtain the final pseudo image output.
[0110] Since random noise is added to weaken the content image features and the underwater noise environment is simulated with the help of this noise, the impact of noise on the synthetic image is inevitable and will reduce the visual effect of the synthetic image. It is difficult or unreliable to obtain a clean reference real image from the synthesized pseudo image, so it is impossible to use the clean image as the target to enhance the noisy image. We propose a fast guided filtering enhancement module that can balance the noise in the generated image and improve the stylization effect. The fast guided filtering enhancement module uses the local linear model and the input image itself as the guide image to guide the enhancement of the initial pseudo sonar underwater dam crack image as the filtered input image. Filtering or edge preservation processing is performed under different variances to generate a pseudo image that is more similar to the real underwater sonar dam crack image. The sonar image dataset is used as a training sample to train the underwater dam crack recognition model, improve the recognition accuracy of dam cracks, and significantly reduce the need for labeled underwater dam crack samples. The overall framework diagram of the pseudo sonar underwater dam crack image generation method based on style transfer;
[0111] like Figure 7The results of the method in this paper are shown. This paper selects optical crack images according to the different degrees of cracking on the building surface, such as horizontal cracks, longitudinal cracks, multiple cracks, collapse, surface protrusions, etc. The first and fourth columns are optical content images, among which the content images of the first image in the first column and the third image in the fourth column are from the crack detection dataset; the content images of the second image in the first column and the first and second images in the fourth column are from the Concrete Crack dataset, and the content image of the third image in the first column is from the pavement crack dataset. The side scanning sonar dataset uses public data. In order to better demonstrate the effect of pseudo image generation under different styles, this paper selects shipwrecks as style images, such as the style images of the first image in the second column and the first image in the fifth column; selects seabed ore as style images, such as the style images of the second image in the fifth column; and seabed topography, such as the style images of the third image in the second column and the third image in the fifth column. From the generated results, we can see that there are only the same target cracks between the generated image and the content image. At the same time, it only has the sonar image style of the style image, and does not contain the target in the sonar image. After being weakened by the SDFWMCI module, specific domain features such as the color in the original content image and the texture on the cracks have been eliminated or faded, which can make the generated image more similar to the real sonar image.
[0112] The generated images of this method can not only expand the training samples of the sub-class recognition model and improve the recognition effect, but also avoid the influence of object color on the stylization effect and the inspection accuracy of dam cracks. Pseudo-sonar images are added with noise to simulate the influence of underwater noise environment on sonar imaging. Figure 7 Taking the first and third generated images in the third column as an example, after being enhanced by the FGFE module, the generated image has a noise shadow similar to the style image. From the results of the first image in the third column and the second image in the sixth column, it can be seen that the introduction of the SCPE module makes the style rendering effect of the generated image more reasonable, and similar positions have the same stylized effect. From the third generated result in the third column, it can be seen that the generated image does not have the material texture of the optical crack target, and the shadow distribution formed by the noise is reasonable. In summary, the pseudo image generated by the pseudo image generation method in this paper is more similar to the real sonar image and has better quality.
[0113] The above shows and describes the basic principles, main features and advantages of the present invention. It should be understood by those skilled in the art that the present invention is not limited to the above embodiments. The above embodiments and descriptions are only for explaining the principles of the present invention. Without departing from the spirit and scope of the present invention, the present invention may have various changes and improvements, which fall within the scope of the present invention. The scope of protection of the present invention is defined by the attached claims and their equivalents.
Claims
1. A pseudo-sonar underwater dam crack image generation method based on style transfer, characterized by: The steps include: Step 1: Construct an underwater dam crack pseudo image generation network, which includes a content image specific domain feature weakening module, an image patch-based style transfer pseudo image generation network, and a fast guided filtering enhancement module; Step 2: Send the optical crack image as the content image data set to the underwater dam body crack pseudo image generation network; the underwater dam body crack pseudo image generation network processes the content image data set as follows: Step 2.1: Use the content image specific domain feature weakening module to weaken the non-specific domain features of the content image dataset to obtain the optical content image dataset after feature weakening; Step 2.2: Use the side-scan sonar image as a style image dataset, and input the style image dataset and the attenuated optical content image dataset into an image patch-based style transfer pseudo image generation network, and obtain the content image patch sequence of the optical content image dataset and the style image patch sequence of the style image dataset through the image patch-based style transfer pseudo image generation network; Step 2.3: Reconstruct the initial pseudo sonar underwater dam crack image using the content image patch sequence and the style image patch sequence; Step 2.4: The initial pseudo-sonar underwater dam crack image is used as the guide image and the filtered input image, and the fast guided filter enhancement module is used to perform noise reduction and enhancement; Step 3: Generate the final pseudo-sonar underwater dam crack image.
2. The method for generating a pseudo-sonar underwater dam crack image based on style conversion according to claim 1 is characterized in that: The process of obtaining the content image patch sequence is: Step 2.2.1: Segment the optical crack image in the optical content image dataset into several image patches through PatchEmbed; Step 2.2.2: The image block is linearly projected and SCPE module is used to obtain a content image patch sequence with position coding; wherein, firstly, a convolution layer and an adaptive average pooling layer are used to perform an average pooling operation on the content image of the image block; the adaptive average pooling layer pools the input feature map into a fixed-size output; the pooled feature map is linearly projected through a convolution layer, and the convolution layer maps the input feature map from 512 channels to 512 channels to generate a content-related position coding; the generated content position coding is adjusted to the same height and width as the style image through bilinear interpolation; finally, the position coding is added to the content features of the content image as the position information of the content features, and a content image patch sequence combined with the position coding is obtained; At the same time, the process of obtaining the style image patch sequence is: Step 2.2.3: Use PatchEmbed to split the side-scan sonar image in the style image dataset into several image patches; Step 2.24: The image block is linearly projected to obtain a style image patch sequence.
3. The method for generating a pseudo-sonar underwater dam crack image based on style conversion according to claim 2 is characterized in that: Image Blocks Semantic content-adaptive positional encoding for: ; ; in, is the average pooling function, is a learnable position encoding function Convolution operation, is a sequence of content image patches A learnable positional encoding of , is the interpolation weight and N is the number of neighboring blocks.
4. The method for generating a pseudo-sonar underwater dam crack image based on style conversion according to claim 1 is characterized in that: Step 2.3 is as follows: Step 2.3.1: The content image patch sequence and the style image patch sequence are fed to the content image encoder and the style image encoder respectively; Step 2.3.2: The outputs of the content image encoder and the style image encoder are fed into a multi-layer decoder; Step 2.3.3: The multi-layer decoder translates the encoded content image patch sequence in a regression manner according to the encoded style image patch sequence, and obtains the initial pseudo sonar underwater dam crack image.
5. The method for generating a pseudo-sonar underwater dam crack image based on style conversion according to claim 4 is characterized in that: The processing of the multi-layer decoder is: The input to the multi-layer decoder consists of a sequence of encoded content image patches and a sequence of style image patches , generate queries (Q) using a sequence of content image patches, and generate keys (K) and values (V) using a sequence of style image patches: ; The calculation process of the output sequence G of the multi-layer decoder is: ; ; ; Among them, the output sequence G of the multi-layer decoder is The shape of the output sequence is not directly upsampled to construct the final result, but a 3-layer CNN decoder is used to refine the subsequent output of the transformer decoder; for each layer, a series of operations are used to expand the scale, including ; This operation converts the output of the multi-layer decoder into a format suitable for convolution operations; the first layer of convolution equals the number of input channels to the embedding dimension and halves the number of output channels, gradually reducing the complexity of the feature space, and using The convolution kernel captures local area features and keeps the output resolution unchanged; the second convolution layer receives the output of the previous layer, further reducing the number of channels and maintaining the same resolution; the last layer reconstructs the resolution Initial pseudo-sonar underwater dam crack image.
6. The method for generating a pseudo-sonar underwater dam crack image based on style conversion according to claim 1, characterized in that: The content image specific domain feature weakening module includes an encoder, a decoder, random noise, a feature whitening module, a Haar wavelet decomposition module and an image reconstruction module, the encoder is VGGEncoder; the decoder is VGGDecoder; the feature whitening module performs whitening transformation in the feature space and redistributes the features through principal component analysis; the Haar wavelet decomposition module decomposes the features into different frequency components for analysis or reconstruction, and obtains low-frequency components and high-frequency components by performing Haar wavelet decomposition, which are passed to the decoder to generate low-frequency feature maps and high-frequency feature maps respectively; the image reconstruction module reconstructs the input image after whitening transformation and Haar wavelet decomposition to generate an image for the features; at the same time, the final results of the whitening operation, noise processing and feature transformation are combined with pixel reconstruction loss and feature loss to achieve reconstruction and generation of feature weakened images.
7. The method for generating a pseudo-sonar underwater dam crack image based on style conversion according to claim 6 is characterized in that: Step 2.1 is described by the formula: ; in, Represents the input image The predicted output of It is a feature extraction encoder that extracts deep features from the input image; It is the function responsible for feature separation, which removes the texture and color features in the image and retains the domain-invariant features; It is a classifier that realizes the mapping between deep features and predictions; A set of instances representing content images; A collection of instances representing sonar crack images; Then use the pre-trained autoencoder to extract the input image features, and reconstruct the input image using pixel reconstruction loss and feature loss; ; in, The total loss for feature-weakened image reconstruction; It is a measure of the difference between the reconstructed image and the input image, which is the pixel reconstruction loss; It is a measure of the difference between the deep features of the reconstructed image and the input image, which is the feature loss; represents the reconstructed output, represents the input image; It is an encoder that extracts deep features of the image; Then perform whitening transformation: ; in, The characteristic value is The diagonal matrix of the covariance matrix, is the orthogonal matrix corresponding to the eigenvectors, represents the whitened content image features, is the content image depth feature; Then, the style features extracted by the encoder from the image are fused with the whitened content features through coloring transformation, expressed as formula (4): ; in, Indicates the dyeing characteristics. The characteristic value is The diagonal matrix of the covariance matrix, is the orthogonal matrix corresponding to the eigenvectors, is the style image depth feature, and finally the merged depth feature Decoded by the decoder to reconstruct the image .
8. The method for generating a pseudo-sonar underwater dam crack image based on style conversion according to claim 1, characterized in that: Step 2.4 is described by the formula: ; in, represents the index of the pixel in the guided filter; The output image is pixel value; is the first pixel value; represents a local window with radius r, Display Window The index of , Display Window The linear coefficients in; given an image Under the input condition, the above formula (22) can be obtained: ; ; in, , To guide the image In the window The variance and mean of represents the regularization parameter related to the degree of smoothness; Since each window is calculated , , but each pixel is contained in multiple windows, and for each pixel, multiple , , multiple , The calculated results are averaged to get the output value. The above process is described as follows: ; ; in, , represents the average coefficient of all windows, and the window is As the center; reduce the number of pixels by downsampling and calculate , The average value is then upsampled to restore to the original size; in high variance areas, The value of is large, the output image It depends on the guide image , thereby retaining edge information; while in low variance areas, The value of is small, the output image Then smoothing is performed; the final pseudo-sonar underwater dam crack image is the output image .