Seabed sediment image data amplification method based on self-attention generative adversarial network

By using a self-attention generation adversarial network in the amplification of seabed bottom image data, the generator network model is simplified and the self-attention mechanism is introduced, the problem of sparse samples of seabed bottom sonar image data is solved, the image quality and model learning ability are improved, and the effective amplification of seabed bottom image data and the classification accuracy are improved.

CN120047805APending Publication Date: 2025-05-27HARBIN UNIV OF SCI & TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510056188.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-14
Publication Date
2025-05-27

AI Technical Summary

Technical Problem

The sparse sample of the subsea subsea subsea subsea subsea subsea subsea subsea subsea subsea subsea subsea subsea subsea subsea subsea subsea subsea subsea. Especially when using deep learning algorithms, the lack of training samples directly leads to a decline in the performance of the classification model.

Method used

The seabed bottom image data amplification method based on self-attention generation adversarial network is adopted. By simplifying the model of the generator network, the self-attention mechanism is introduced, a generator network with strong feature learning ability is built, and a discriminator network with strong recognition ability is built to achieve effective amplification of seabed bottom image data.

Benefits of technology

The quality of generated images is improved, the model's learning ability of seabed bottom image features is enhanced, and the performance bottleneck caused by sparse samples is overcome. The generated seabed bottom image quality is better, which can effectively alleviate the problem of low classification accuracy of seabed bottom image.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120047805A_ABST
    Figure CN120047805A_ABST
Patent Text Reader

Abstract

The invention discloses a seabed sediment image data amplification method based on a self-attention generative adversarial network, and the method comprises the steps: firstly constructing a generator network model which is high in feature learning capability and is simple in structure through simplifying a model of a generator network and introducing a self-attention mechanism, so as to learn the real distribution of a real seabed sediment acoustic image; then, a discriminator network based on a self-attention convolutional neural network and having strong recognition capability is constructed by introducing a self-attention mechanism, and the data source distinguishing capability of the discriminator network is improved; and finally, realizing alternate evolution of the discriminator network and the generator network based on a binary minimum-maximum game strategy, generating a simulation image similar to a real seabed sediment acoustic image, and realizing seabed sediment image data amplification. According to the method, effective seabed sediment image data amplification is realized, and the problem of sparse seabed sediment sonar image data samples is solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of seabed sediment image processing, and relates to a method for augmenting seabed sediment image data, specifically to a method for augmenting seabed sediment image data based on a self-attention generative adversarial network. Background Art

[0002] In view of the current situation of the gradual depletion of land resources and the continuous intensification of population pressure, the pace of the social and economic development and scientific and technological progress of human society is facing the challenge of slowing down. Against this background, the ocean, as the most extensive natural ecosystem on the earth, contains rich and diverse production and living resources. Therefore, ocean development has become a key strategy to address the above problems. With its huge potential and diverse characteristics, ocean resources can supply indispensable raw materials for many fields such as industry and agriculture. Moreover, ocean development activities have become an important force and core driving force for promoting scientific and technological progress. With the frequent intensification of human maritime activities, accurately obtaining ocean environmental information has become an indispensable prerequisite for ocean development and utilization. Seabed sediment, which is composed of the ocean floor and the substances deposited on it through long-term ocean movements, constitutes a key part of ocean environmental information. The application of seabed sediment classification technology shows its importance in several key dimensions: First, accurately obtaining seabed sediment information can plan safe ship navigation routes, effectively avoid navigation risks, and provide a solid basis and guarantee for navigation safety. Second, the mastery of sediment information is decisive for clarifying the distribution of seabed resources, providing detailed data bases and guidance for large-scale projects such as underwater optical fiber laying, port construction, seabed metal mineral mining, seabed oil and gas resource development, and cross-sea bridges. Third, accurately obtaining sediment information can also effectively monitor the ocean pollution situation, track the pollution source and its diffusion path, and provide an efficient and accurate monitoring tool for national environmental comprehensive management and ecological protection. Therefore, accurately obtaining seabed sediment information shows crucial value in many fields such as environmental monitoring, resource development and utilization, and scientific research, and is a key factor in promoting the development of related fields.

[0003] Seabed sediment images, as low-resolution grayscale images, are mainly composed of grayscale information and sediment texture features. Compared with optical images, the amount of information they contain is usually limited. When collecting seabed sediment data, due to the interference of natural factors such as waves and currents, the amount of seabed sediment data obtained is often small, resulting in the problem of sparse seabed sediment sonar image data samples. This sparsity poses a challenge to the training of seabed sediment classification models, especially when using deep learning algorithms. Due to their strong dependence on data, the lack of training samples will directly lead to a decline in the performance of the classification model. In view of the above situation, in order to improve the accuracy of the seabed sediment classification model, it is necessary to introduce a large amount of training data. By increasing the number of data samples, the model's ability to learn various seabed sediment image features can be effectively enhanced, thereby overcoming the performance bottleneck caused by sparse samples.

[0004] Traditional generative adversarial networks tend to adopt a fully connected structure when constructing a generator network. Although this design may enhance the generator's ability to extract features to a certain extent, it also significantly increases the complexity of the network, thereby increasing the difficulty of training and the consumption of computing resources. More importantly, when faced with sparse sample data, the fully connected structure may lead to a decline in model performance, because complex network structures are more likely to overfit under limited data. On the other hand, the generator network in the image generation task generally relies on convolutional layers to build. However, the operation of the convolutional layer mainly focuses on the extraction of local neighborhood features of the image, which limits the model's ability to learn the global data distribution of the image and makes it difficult to fully capture the global feature information of the image. This limitation may prevent the generator from generating images with global consistency and high-quality details. In addition, the performance of the discriminator network is also crucial to the training process of the generative adversarial network and the evolution of the generator network. If the discriminator's recognition ability is insufficient, it will directly affect the learning direction and efficiency of the generator, which may lead to instability in the training process of the generative adversarial network, and then affect the convergence of the generator network and the quality and clarity of the final generated image.

[0005] Therefore, in order to build a generative adversarial network that can generate clear images with stable quality, the key is to design a generator structure that can effectively extract global features and control the complexity of the network, and to build a discriminator network with strong recognition ability to ensure the stability of the generative adversarial network training and the good convergence of the generator network. Summary of the invention

[0006] To solve the problem of insufficient sonar image data samples of seabed sediment and improve the quality of images generated by the generative adversarial network, the present invention uses a small number of collected sonar images of seabed sediment and proposes a method for augmenting seabed sediment image data based on a self-attention generative adversarial network model. By simplifying the model construction of the generator network, the method constructs a generator network with a simple structure, low model complexity, and strong feature learning ability. At the same time, it utilizes the advantage of the self-attention mechanism to capture global image information to improve the model's representation ability, enhance the quality of the generated images, achieve effective augmentation of seabed sediment image data, and solve the problem of sparse sonar image data samples of seabed sediment.

[0007] The object of the present invention is achieved by the following technical solutions:

[0008] A method for augmenting seabed sediment image data based on a self-attention generative adversarial network includes the following steps:

[0009] Step 1: Generate a set of random noise data X that conforms to the Gaussian distribution rn ;

[0010] Step 2: Randomly sample the generated set of random noise data X rn to obtain a sampled dataset of random noise samples X ns ;

[0011] Step 3: Randomly sample real seabed sediment images to obtain a sampled dataset of real seabed sediment images X r ;

[0012] Step 4: Determine the basic structure of the generative adversarial network, and design the internal parameters of the network according to the dimension of the input noise and the size of the real seabed sediment image to construct the generative adversarial network;

[0013] Step 5: Input the sampled dataset of random noise samples X ns into the generator network, and the generator network maps the sampled dataset of random noise samples X ns to a dataset of seabed sediment images with the same size as the real seabed sediment images X g ;

[0014] Step 6: Input the generated dataset of seabed sediment images X g and the sampled dataset of real seabed sediment images X r into the discriminator network respectively, and the discriminator network determines the source of the input data, that is, whether the input data comes from the generated dataset of seabed sediment images X g or the sampled dataset of real seabed sediment images X r ;

[0015] Step 7: Calculate the difference between the generated seabed sediment image and the real seabed sediment image according to the loss function and optimization strategy of the discriminator network, and adjust the parameters of the discriminator network;

[0016] Step 8: Adjust the parameters of the generator network according to the output result of the discriminator network, as well as the optimization strategy and loss function of the generator;

[0017] Step 9: Use the adjusted generative adversarial network to batch generate seabed sediment images, and repeat Steps 1-8 until the discriminator network can no longer distinguish the generated seabed sediment image dataset X g and the real seabed sediment image dataset X r between the differences or the predetermined number of training times is completed;

[0018] Step 10: Use the trained generative adversarial network to batch generate qualified seabed sediment image data to achieve effective amplification of seabed sediment image data.

[0019] Compared with the prior art, the present invention has the following advantages:

[0020] (1) Simplify the model of the generator network by pruning the fully connected layer in the generator network, making the training of the generator network easier and reducing the consumption of computing resources;

[0021] (2) By introducing the self-attention mechanism, a generator network model with strong feature learning ability and capable of obtaining global image features is constructed, overcoming the problems of poor feature extraction ability of traditional models based on convolutional layers and difficulty in obtaining global image information, and improving the ability of the generator network to learn the true distribution of real seabed sediment acoustic images;

[0022] (3) By introducing the self-attention mechanism, a discriminator network based on self-attention convolutional neural network with strong recognition ability is constructed, improving the ability of the discriminator network to distinguish whether specific data comes from the real seabed sediment acoustic image dataset or the generated dataset, and thus promoting the evolution of the generator network.

[0023] (4) It can generate seabed sediment images with better image quality, thereby alleviating the problem of low classification accuracy of seabed sediment images caused by sparse image samples. Description of the Drawings

[0024] Figure 1 is a flowchart of the method for amplifying seabed sediment image data based on the self-attention generative adversarial network of the present invention.

[0025] Figure 2 is a structural diagram of the self-attention generative adversarial network proposed by the present invention.

[0026] Figure 3 There are three types of real seabed sediment images.

[0027] Figure 4 There are three types of seabed sediment images generated by the method proposed in the present invention. Detailed implementation manners

[0028] The technical solution of the present invention will be further described below in conjunction with the accompanying drawings, but it is not limited thereto. Any modification or equivalent replacement of the technical solution of the present invention without departing from the spirit and scope of the technical solution of the present invention shall be covered by the protection scope of the present invention.

[0029] Traditional data generation methods require constructing an accurate mathematical expression relationship between input data and output data, which greatly increases the difficulty of data augmentation. Different from traditional data generation methods, generative adversarial networks can automatically construct a mapping relationship between input data and output data by learning the patterns between input and output data, and have good effects in the field of image processing research. There is a fully connected layer structure in traditional generative adversarial networks. When facing sparse sample data, the fully connected structure may lead to a decline in model performance because complex network structures are more likely to overfit under limited data. In addition, the generator network in image generation tasks generally relies on convolutional layers to construct. However, the operations of convolutional layers mainly focus on the extraction of local neighborhood features of images, which limits the model's ability to learn the global data distribution of images and is difficult to comprehensively capture the global feature information of images. In view of this, the present invention provides a method for augmenting seabed sediment image data based on a self-attention generative adversarial network. First, a generator network model with strong feature learning ability and simple structure is constructed by simplifying the model of the generator network and introducing a self-attention mechanism to learn the real distribution of real seabed sediment acoustic images; then, a discriminator network based on a self-attention convolutional neural network with strong recognition ability is constructed by introducing a self-attention mechanism to improve the discriminator network's ability to distinguish the source of data; finally, based on the binary min-max game strategy, the discriminator network and the generator network are alternately evolved to generate simulated images similar to real seabed sediment acoustic images, realizing the augmentation of seabed sediment image data. As Figure 1 , the specific steps are as follows:

[0030] Step 1: Generate a set X of random noise data that conforms to the Gaussian distribution rn .

[0031] In this step, the dimension of the random noise is not limited. However, through multiple experimental verifications, it is found that when the size of the random noise data that conforms to the Gaussian distribution is 1×1×100, the image quality generated by the generative adversarial network is better. In addition, to ensure the quality of the finally generated images, the set X of random noise data is increased as much as possible within the limits of computer memory and performancern Sample size.

[0032] Step 2: Randomly sample the generated random noise data set X rn to obtain the sampled random noise sample data set X ns .

[0033] In this step, through a large number of experimental verifications, when the sample size of each batch of random noise sample data set X ns is 2 to the power of n, the quality of the generated seabed sediment image samples is better. For the convenience of constructing the model of the generative adversarial network, the sample size of each batch of random noise sample data set X ns is determined to be 64.

[0034] Step 3: Randomly sample real seabed sediment images to obtain the sampled real seabed sediment image data set X r .

[0035] In this step, to ensure the smooth progress of the model training process of the generative adversarial network, it is necessary to ensure that the sample size of the real seabed sediment image data set X r is the same as that of the random noise sample data set X ns . Here, the real seabed sediment image can be a side-scan sonar image, a multi-beam sonar image, a synthetic aperture sonar image, an underwater optical image, etc.

[0036] Step 4: Determine the basic structure of the generative adversarial network, design the internal parameters of the network according to the dimension of the input noise and the size of the real seabed sediment image, and construct the generative adversarial network.

[0037] In this step, as Figure 2 shown, the constructed generator network uses a total of 5 transposed convolutional layers and 2 self-attention layers, where: Transposed Convolutional Layer 1 and Transposed Convolutional Layer 3 are both followed by a self-attention layer. The transposed convolutional layer mainly performs upsampling operations, mapping the low-resolution feature map to the high-resolution space, thereby generating more detailed image details, and then mapping the noise data into seabed sediment images. The self-attention layer calculates the correlation between features at different positions through the self-attention mechanism to generate an attention feature map. The generator network uses this attention feature map to weight the features at different positions, so that when generating each pixel, it can comprehensively consider global information. This process not only improves the detail quality of image generation but also enhances the overall coherence and consistency of the image. The overall parameters of the generator network are shown in Table 1. As Figure 2As shown, the constructed discriminator network uses a total of 4 convolutional layers, 2 self-attention layers, and 1 fully connected layer, where: A self-attention layer is connected after both convolutional layer 1 and convolutional layer 3, and a fully connected layer is connected after convolutional layer 4. The convolutional layers capture local features such as edges and textures in the input image through convolution operations, which are the basis for subsequent high-level feature extraction and classification. In addition, through the sliding of the convolution kernel and the generation of feature maps, the convolutional layers achieve dimensionality reduction of the input data, which not only reduces the computational amount of subsequent layers but also helps to extract more compact and representative features. The self-attention layer can model long-range dependencies in the image and construct differential features between real images and generated images. The discriminator can utilize this feature to check whether the detailed features of more distant parts in the image are consistent. This helps the discriminator more finely distinguish the differences between real images and generated images, thereby improving the discrimination accuracy. The fully connected layer is responsible for integrating the features extracted by the convolutional layers and the self-attention layer, combining local features and global context information to form a more comprehensive feature representation, and outputting the probability value of whether the input image is real, thereby achieving classification decisions. The overall parameters of the discriminator network are shown in Table 2. The mathematical expression of the self-attention mechanism is where Q is the query, K is the key, V is the value, and d isa represents the dimension of the query Q and the key K, and X isa is the input, and K T is the transpose of the key K.

[0038] In addition, during the training process of the generative adversarial network, the number of training times, the initial learning rate, and the batch size are determined to be 5000, 1e -4 respectively, and the RMSprop optimization strategy is used in the training to guide the network to update the parameters.

[0039] Table 1 Overall parameters of the generator network

[0040]

[0041] Table 2 Overall parameters of the discriminator network

[0042]

[0043] Step 5: Input the random noise sample dataset X ns into the generator network. The generator network maps the random noise sample dataset X ns to a seabed sediment image dataset X g with the same size as the real seabed sediment image according to the current generator network parameters.

[0044] Step 6: The generated seabed sediment image dataset X g and the real seabed sediment image dataset X rInput into the discriminator network respectively, and the discriminator network determines the source of the input data, that is, whether the input data comes from the generated seabed sediment image dataset X g or the real seabed sediment image dataset X r .

[0045] Step 7: Calculate the difference between the generated seabed sediment image and the real seabed sediment image according to the loss function and optimization strategy of the discriminator network, and adjust the parameters of the discriminator network to improve the ability of the discriminator network to identify the data source.

[0046] In this step, in order to ensure that the discriminator network can effectively learn the difference between the generated seabed sediment image and the real seabed sediment image, the loss function of the discriminator network is determined as where D(·) represents the discriminator network, G(·) represents the generator network, z represents the input signal, P g represents the data distribution of the generated image, x represents the real seabed sediment image data, P r represents the data distribution of the real seabed sediment image, and the RMSprop optimization strategy is used to guide the discriminator network to update the parameters. In addition, the number of training times, the initial learning rate, and the batch size of the discriminator network are determined to be 5000, 1e -4 、64 respectively.

[0047] Step 8: Adjust the parameters of the generator network according to the output result of the discriminator network, the optimization strategy and the loss function of the generator, so that the difference between the seabed sediment image generated by the generator network and the real seabed sediment image gradually decreases.

[0048] In this step, in order to ensure that the generator network can effectively learn the mapping relationship between the randomly generated noise data that conforms to the Gaussian distribution and the real seabed sediment image, the loss function of the generator network is determined as where D(·) represents the discriminator network, G(·) represents the generator network, z represents the input signal, P g represents the data distribution of the generated image, and the RMSprop optimization strategy is used to guide the process of adjusting the parameters of the generator network. In addition, the number of training times, the initial learning rate, and the batch size of the generator network are determined to be 5000, 1e -4 、64 respectively.

[0049] Step 9: Use the adjusted generative adversarial network to batch generate seabed sediment images, and repeat steps 1-8 until the discriminator network cannot identify the difference between the generated seabed sediment image dataset X g and the real seabed sediment image dataset X r or the predetermined number of training times is completed.

[0050] Step 10: Use the trained generative adversarial network to batch generate qualified seabed sediment image data to achieve effective amplification of seabed sediment image data.

[0051] The real three different types of seabed sediment images are as Figure 3 shown, and the generated three different types of seabed sediment images are as Figure 4 shown. From Figure 3 and Figure 4 it can be seen that the seabed sediment images generated by the method of the present invention are relatively clear, the features of different types of seabed sediment images are obvious, and the visual difference between the generated seabed sediment images and the real seabed sediment images is very small, which proves the effectiveness of the present invention.

Claims

1. A method for amplifying seabed sediment image data based on a self-attention generative adversarial network, characterized in that The method comprises the following steps: Step 1: Generate a random noise data set X that conforms to the Gaussian distribution rn ; Step 2: Generate random noise data set X rn Perform random sampling to obtain the sampled random noise sample data set X ns ; Step 3: Randomly sample real seabed images to obtain the sampled real seabed image dataset X r ; Step 4: Determine the basic structure of the generative adversarial network, design the internal parameters of the network according to the dimension of the input noise and the size of the real seabed sediment image, and construct the generative adversarial network; Step 5: Substitute the random noise sample data set X ns Input the generator network, and the generator network converts the random noise sample data set X into ns The seabed image dataset X is mapped to the same size as the real seabed image g ; Step 6: Generate the generated seafloor image dataset X g and the real seabed image dataset X r The discriminator network determines the source of the input data, that is, the input data comes from the generated seabed bottom image dataset X g Or a real seabed image dataset X r ; Step 7: Calculate the difference between the generated seabed sediment image and the real seabed sediment image according to the loss function and optimization strategy of the discriminator network and adjust the parameters of the discriminator network; Step 8: Adjust the parameters of the generator network according to the output of the discriminator network and the optimization strategy and loss function of the generator; Step 9: Use the adjusted generative adversarial network to batch generate seabed sediment images, and repeat steps 1 to 8 until the discriminator network cannot recognize the generated seabed sediment image dataset X g and the real seabed image dataset X r The difference between or the completion of the scheduled number of training sessions; Step 10: Use the trained generative adversarial network to batch generate qualified seabed sediment image data to achieve effective seabed sediment image data amplification.

2. The method for amplifying seabed sediment image data based on a self-attention generative adversarial network according to claim 1 is characterized in that In step 2, each batch of random noise sample data set X ns The sample size is 2 to the power of n.

3. The method for amplifying seabed sediment image data based on a self-attention generative adversarial network according to claim 1, characterized in that In step 3, the real seabed bottom image dataset X r The sample size and random noise sample data set X ns The sample sizes are the same.

4. The method for amplifying seabed sediment image data based on a self-attention generative adversarial network according to claim 1, characterized in that In step 3, the real seabed sediment image is a side scan sonar image, a multi-beam sonar image, a synthetic aperture sonar image, or an underwater optical image.

5. The method for amplifying seabed sediment image data based on a self-attention generative adversarial network according to claim 1, characterized in that In step 4, the generator network of the generative adversarial network includes 5 transposed convolution layers and 2 self-attention layers, wherein: transposed convolution layer 1 and transposed convolution layer 3 are each followed by a self-attention layer, the transposed convolution layer performs an upsampling operation to map the low-resolution feature map to the high-resolution space, thereby generating more detailed image details, and then mapping the noise data to the seabed bottom image; the self-attention layer calculates the correlation between features at different positions through the self-attention mechanism to generate an attention feature map, and the generator network uses the attention feature map to weight features at different positions, so that global information can be comprehensively considered when generating each pixel; the discriminator network of the generative adversarial network includes 4 convolutional layers, 2 self-attention layers and 1 fully connected layer, among which: convolutional layer 1 and convolutional layer 3 are each followed by a self-attention layer, and convolutional layer 4 is followed by a fully connected layer. The convolutional layer captures the local features in the input image through convolution operations, and realizes dimensionality reduction processing of the input data through sliding of the convolution kernel and generation of feature maps; the self-attention layer models the long-distance dependency in the image and constructs the difference features between the real image and the generated image; the fully connected layer is responsible for integrating the features extracted by the convolutional layer and the self-attention layer, combining the local features with the global context information to form a more comprehensive feature representation, and outputs the probability value of whether the input image is real, thereby realizing classification decisions.

6. The method for amplifying seabed sediment image data based on a self-attention generative adversarial network according to claim 1, characterized in that In step 7, the loss function of the discriminator network is Where D(·) represents the discriminator network, G(·) represents the generator network, z represents the input signal, and P g represents the data distribution of the generated image, x represents the real seabed bottom image data, P r It represents the real distribution of seabed image data and uses the RMSprop optimization strategy to guide the discriminator network to update its parameters.

7. The method for amplifying seabed sediment image data based on a self-attention generative adversarial network according to claim 1, characterized in that In step 8, the loss function of the generator network is Where D(·) represents the discriminator network, G(·) represents the generator network, z represents the input signal, and P g Represents the data distribution of generated images and uses the RMSprop optimization strategy to guide the parameter adjustment process of the generator network.