An underwater video enhancement method based on CycleGAN
Through the CycleGAN-based underwater video enhancement method, unsupervised training is used to use information interaction between video frames to solve problems such as color distortion and texture blur in underwater videos, and efficient video enhancement effect is achieved.
Patent Information
- Application Number
- CN202411668015.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-21
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2044-11-21
Smart Images

Figure CN119152359B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of video enhancement technology, and in particular relates to an underwater video enhancement method based on CycleGAN. Background Art
[0002] Video enhancement involves processing and improving raw video using a range of techniques and algorithms to enhance image quality, increase detail, reduce noise, and improve visual quality. For example, a series of techniques selectively highlight valuable information while suppressing unnecessary information, making the enhanced video more suitable for human or machine analysis and processing. Video enhancement techniques include adjusting resolution, increasing bitrate, reducing noise, and adjusting brightness, contrast, color, and other parameters to achieve clearer images, more vivid colors, and richer details.
[0003] When conducting underwater exploration missions, the complex underwater environment causes light propagation to be affected by absorption and scattering by the water medium and by impurities in the water. This can lead to problems such as color distortion, blurred textures, low contrast, and uneven lighting in underwater videos. With the rapid development of computer technology and image processing techniques, underwater video enhancement technology has also made significant progress.
[0004] In terms of specific technical implementation, early research focused primarily on hardware devices, such as the polarizer method, which uses the principle of polarizers to remove scattered light. Later, with the development of computer technology, researchers began to use computer algorithms for underwater video enhancement. A typical method is the fast shutter threshold imaging method, which uses an imaging device with an ultra-high shutter speed to rapidly open and close the shutter to remove scattered light. With the development of technology, a variety of enhancement techniques and methods have been proposed and applied. Traditional enhancement methods can be divided into two types based on the scope of enhancement: spatial domain enhancement and frequency domain enhancement. Spatial domain enhancement directly processes each pixel in the image; for example, methods such as histogram equalization, grayscale conversion, and contrast stretching. Frequency domain enhancement processes the spectral components of the image after Fourier transform, and then obtains the desired image through inverse Fourier transform. By selecting appropriate filters, image noise can be removed, edges can be highlighted, and the visual quality of the image can be improved.
[0005] For underwater data, researchers have proposed a variety of algorithms to address these challenges. There are numerous underwater image enhancement and restoration methods, broadly categorized as non-physical model methods and physical model-based methods. Non-physical model methods, which fall under the image enhancement category, do not rely on mathematical models of underwater optical imaging and improve visual quality by adjusting the image's pixel values. Physical model-based methods mathematically model the degradation process of underwater images, estimate model parameters, and invert the degradation process to obtain clear underwater images. These methods fall under the image restoration category.
[0006] The recent introduction of deep learning methods has brought new possibilities for underwater video enhancement. For example, deep learning models such as convolutional neural networks (CNNs), transformers, and generative adversarial networks (GANs) can automatically learn the characteristics of underwater videos and optimize model parameters using large amounts of training data, thereby achieving better video enhancement results. The basic structure of a CNN consists of an input layer, convolutional layers, pooling layers, activation layers, and fully connected layers. In image-processing CNNs, the input layer typically represents an image's pixel matrix. The convolutional layers are the core of the CNN, extracting image features through convolution operations. Pooling layers further reduce the spatial size of the data, reducing the number of parameters and preventing overfitting. Activation layers introduce nonlinearity, enabling the neural network to learn complex image features. The Transformer architecture consists of an encoder and a decoder. The Vision Transformer (ViT), proposed in 2020 for image processing, divides the image into 16x16 blocks, making the Transformer more suitable for image processing tasks. The Generative Adversarial Network (GAN) was proposed by Goodfellow et al. in 2014. It consists of two core components: a generator and a discriminator, which compete with each other to learn the data distribution. The generator's task is to learn from random noise and generate data similar to real data, while the discriminator's task is to distinguish between generated data and real data. This competitive relationship drives the model to continuously evolve, making the generated data gradually closer to the real data distribution.
[0007] Based on the development of deep learning technology, Christian Ledig et al. proposed a Generative Advancement (GAN) model in 2016 to enhance the resolution of images and videos. This model uses SRResNet as a generative network for super-resolution, known as the SRGAN model. This model can produce images with high perceptual quality. They also proposed a new image quality evaluation metric, the mean-opinion-score (MOS), to reflect the degree of similarity between the test object and the real high-resolution image.
[0008] In 2019, Chongyi Li et al. from Tianjin University constructed the Underwater Image Enhancement Benchmark (UIEBDataset), which includes 950 real-world underwater images, 890 of which have corresponding reference images, and the remaining 60 underwater images serve as challenge data. They also proposed Water-Net for underwater image enhancement on the UIEB dataset. The Water-Net model enhances the input image through white balancing (WB), gamma correction (GC), and histogram equalization (HE). It fuses the input with the predicted confidence map to achieve the enhanced result. The input is first transferred to the refined input through the Feature Transformation Unit (FTU), and then the confidence map is predicted. Finally, the enhanced result is obtained by fusing the improved input and the corresponding confidence map.
[0009] With the development of GANs, researchers proposed the CycleGAN generative adversarial network (CycleGAN). In 2023, Qiu Wan and others from Dalian Ocean University proposed an underwater image enhancement algorithm based on CycleGAN super-resolution reconstruction. This algorithm achieves super-resolution reconstruction of underwater images, forming high-resolution underwater images. It also allows for color learning of underwater images from terrestrial images, effectively eliminating color deviation in underwater images and achieving the dual goals of color correction and clarity enhancement for underwater images.
[0010] In 2024, Dazhao Du et al. developed the first large-scale high-resolution underwater video enhancement benchmark (UVEB) to address the differences between underwater images and videos. This benchmark aims to promote the development of underwater video enhancement. They also proposed the UVE-Net, which converts the current frame information into a convolution kernel and transfers it to adjacent frames for efficient inter-frame information exchange. This method fully utilizes the redundant degradation information in underwater videos to better accomplish video enhancement tasks. Summary of the Invention
[0011] To solve the above technical problems, the present invention proposes an underwater video enhancement method based on CycleGAN, which utilizes the information interaction between video frames and does not require a one-to-one correspondence between the source video and the target video, thereby improving the underwater video enhancement effect.
[0012] To achieve the above object, the present invention provides an underwater video enhancement method based on CycleGAN, comprising: obtaining an underwater video dataset, preprocessing underwater video data in the underwater video dataset, and obtaining preprocessed underwater video data;
[0013] The preprocessed underwater video dataset is input into the initial video enhancement model built based on CycleGAN for training to obtain the optimized video enhancement model;
[0014] Target underwater video data is obtained, and the optimized video enhancement model is input to obtain target underwater video enhancement data.
[0015] Optionally, the initial video enhancement model constructed based on CycleGAN includes a generator network and a discriminator network, wherein the generator network includes a first generator and a second generator;
[0016] After the original video frame sequence is input, the intermediate frame is first extracted and input into the first generator. The initial feature extraction is completed through four times downsampling and convolution. The intermediate extracted features are generated through 30 residual blocks. The first convolution kernel is generated through downsampling. At the same time, the first residual result is further downsampled through 30 residual blocks to generate the second convolution kernel.
[0017] In the second generator, the input frame sequence is first downsampled by a factor of four, then subjected to three-dimensional convolution to extract features. After that, each individual frame is convolved separately with the first convolution kernel generated by the first generator using the intermediate frame. The first convolution result is upsampled by a factor of two and then passed through 15 three-dimensional residual blocks. The second output generated by the first generator is then used as the convolution kernel for each individual frame and convolved again to complete the information exchange between frames.
[0018] In the discriminator network, the discriminator divides the input frame sequence into single frames for separate judgment, and averages the results of all frames as the final judgment result of the frame sequence. After the image sequence is input, the image sequence is divided into pictures by frame, and first undergoes an initial convolution and activation layer for preliminary feature extraction, and then inputs the convolution block. The structure of the convolution block is to convolve the input once and then downsample it by one time, and input the activation layer after instance normalization. Three convolution blocks are used continuously in the discriminator network, each convolution block includes one downsampling, and finally three downsampling operations are implemented; the output result is convolved and activated using the sigmoid function, and all frames are averaged after processing, and the result is divided into two to complete the judgment.
[0019] Optionally, the preprocessed underwater video dataset is input into the initial video enhancement model built based on CycleGAN for training. The optimized video enhancement model obtained includes:
[0020] S1. Mapping the pre-processed underwater video data to generate underwater video enhancement data through the first generator;
[0021] S2, discriminating the underwater video enhancement data through the discriminator, and if the underwater video enhancement data is consistent with the pre-processed underwater video data, proceeding to S3; if not, modifying the discriminator parameters and returning to S1;
[0022] S3. Mapping the underwater video enhancement data to generate unenhanced underwater video data through the second generator.
[0023] Optionally, the loss function between the first generator and the discriminator is:
[0024]
[0025] Among them, X, Y represent the source domain and target domain respectively, p data (x) is the data distribution of source domain X, p data (y) is the data distribution of the target domain Y, E y ~p data (y) means y obeys p data (y) to find the expectation, Ex~p data (x) means x obeys p data (x), G represents the generator, D Y represents the discriminator, x is an element of the source domain X, y is an element of the target domain Y, and D Y (y) indicates whether the discriminator identifies y as belonging to the Y domain, G(x) indicates the input x element of the generator G, and D Y (G(x)) represents the discriminator D Y Don't input x into the G generator to see if the result of the input belongs to the Y domain.
[0026] Optionally, the cycle consistency loss between the first generator and the second generator is:
[0027]
[0028] Among them, E y ~p data (y) means y obeys p data (y) to find the expectation, Ex~p data (x) means x obeys p data (x), G is a generator that converts elements of the X domain into the Y domain, F is a generator that converts elements of the Y domain into the X domain, x is an element of the X domain, y is an element of the Y domain, G(x) means using generator G to convert x into an image in the Y domain, F(G(x)) means first using generator G to convert x into an image in the Y domain, and then using generator F to convert the result of generator G back to the X domain, F(y) means using generator F to convert y into an image in the X domain, G(F(y)) means first using generator F to convert y into an image in the X domain, and then using generator G to convert the result of generator F back to the y domain.
[0029] Optionally, the identity loss between the first generator and the second generator is:
[0030]
[0031] Among them, E y ~p data (y) means y obeys p data (y) to find the expectation, Ex~p data (x) means x obeys p data (x), G is a generator that converts elements of the X domain into the Y domain, F is a generator that converts elements of the Y domain into the X domain, x is an element of the X domain, y is an element of the Y domain, G(y) represents the image of y converted to the Y domain using the generator G, and F(x) represents the image of x converted to the X domain using the generator F.
[0032] An electronic device includes: a memory and a processor; the memory is used to store a program; the processor is used to execute the program to implement the underwater video enhancement method based on CycleGAN.
[0033] A readable storage medium stores a computer program, which, when executed by a processor, implements the underwater video enhancement method based on CycleGAN.
[0034] Technical effect of the present invention: The present invention discloses an underwater video enhancement method based on CycleGAN, which uses CycleGAN as a model for unsupervised training without paired data, reducing the model's requirements for the data set while avoiding the mode collapse problem that is prone to CycleGAN. The present invention uses the UEVB large-scale high-resolution data set, has better video enhancement capabilities, and can meet high-resolution input. The network structure in the present invention uses information interaction between frames, uses intermediate frame features as convolution kernels, and passes the information of the auxiliary network to the main network information, which speeds up the training speed while improving the accuracy of the model. The present invention has strong universality. Thanks to the powerful fitting ability of the network, even in the face of underwater conditions under different circumstances, good results can be achieved after perfect training. BRIEF DESCRIPTION OF THE DRAWINGS
[0035] The accompanying drawings, which constitute part of this application, are intended to provide a further understanding of this application. The exemplary embodiments and descriptions of this application are intended to explain this application and do not constitute an improper limitation on this application. In the accompanying drawings:
[0036] Figure 1 Schematic diagram of a process of underwater video enhancement method based on CycleGAN according to an embodiment of the present invention;
[0037] Figure 2 This is a schematic diagram of the network structure of a generator according to an embodiment of the present invention;
[0038] Figure 3 Schematic diagram of the discriminator network structure according to an embodiment of the present invention;
[0039] Figure 4 This is a schematic diagram of cycle consistency loss according to an embodiment of the present invention;
[0040] Figure 5 Schematic diagram of the enhancement effect of different methods on real underwater videos in the embodiments of the present invention, where a is RAW, b is CycleGAN (old), c is FA+Net, d is UVE_Net, e is the present invention, and f is GT. DETAILED DESCRIPTION
[0041] It should be noted that, in the absence of conflict, the embodiments and features of the embodiments in this application can be combined with each other. The present application will be described in detail below with reference to the accompanying drawings and in combination with the embodiments.
[0042] It should be noted that the steps shown in the flowcharts of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions, and that, although a logical order is shown in the flowcharts, in some cases, the steps shown or described can be executed in an order different from that shown here.
[0043] like Figure 1 As shown, this embodiment provides an underwater video enhancement method based on CycleGAN, including: obtaining an underwater video dataset, preprocessing underwater video data in the underwater video dataset, and obtaining preprocessed underwater video data;
[0044] The preprocessed underwater video dataset is input into the initial video enhancement model built based on CycleGAN for training to obtain the optimized video enhancement model;
[0045] Obtain target underwater video data, input the optimized video enhancement model, and obtain target underwater video enhanced data.
[0046] Further, such as Figure 2 The generator network used in this example follows the UVE-Net model's multi-frame input model, with the intermediate frames serving as convolution kernels. The raw image represents a short sequence of extracted images from a real underwater video, while the result represents the enhanced underwater video generated by the generator. mid is the intermediate frame in the video, R is the residual block, and kernel uses the intermediate frame as the convolution kernel. The specific implementation process is as follows: The initial video enhancement model built based on CycleGAN includes a generator network and a discriminator network. The generator network includes a first generator (AE-NET) and a second generator (VE-NET).
[0047] After the original video frame sequence is input, the intermediate frame is first extracted and input into the first generator. After four times downsampling and convolution, preliminary feature extraction is completed, and the intermediate extracted features are generated through 30 residual blocks. The first convolution kernel is generated by downsampling, and the first residual result is downsampled through 30 residual blocks to generate the second convolution kernel; in the second generator, the input frame sequence is first downsampled four times, and then three-dimensional convolution is performed to extract features. After that, convolution is performed separately for each individual frame, and the convolution kernel is the first convolution kernel generated by the intermediate frame in the first generator; after the first convolution result is upsampled twice and then passes through 15 three-dimensional residual blocks, the second output generated in the first generator is used as the convolution kernel of each individual frame for convolution again to complete the information exchange between frames;
[0048] In most cases, the water content and degradation levels between adjacent frames are very similar. Adjacent frames can be subjected to similar feature extraction and enhancement processes to speed up inference. Downsampled frames share similar content and degradation processes with the original frames, which leads to similar enhancement processes. Therefore, the enhancement process of the low-resolution downsampled frame can be leveraged to guide more efficient directional enhancement of the original frame.
[0049] like Figure 3 In the discriminator network shown, the discriminator divides the input frame sequence into single frames for separate judgment, and the results of all frames are averaged as the final judgment result of the frame sequence. After the image sequence is input, the image sequence is divided into pictures by frame. First, it undergoes an initial convolution and activation layer for preliminary feature extraction, and then inputs the convolution block. The structure of the convolution block is to convolve the input once and then downsample it by one. After instance normalization, it is input into the activation layer. Three convolution blocks are used continuously in the discriminator network. Each convolution block contains one downsampling, and finally three downsampling operations are realized; the output result is convolved and activated with a sigmoid function. After all frames are processed, they are averaged and the results are divided into two to complete the judgment. Figure 3 Where f is the number of frames, i.e., the number of images in the input sequence, and IN is instance normalization (InstanceNorm). After the image sequence is input, it is divided into images by frame. After the initial convolution processing, the divided images are sent to the convolution block, which downsamples and extracts features, doubling the sample size each time for a total of three downsamplings. The discriminator in this invention uses the PatchGAN in the original network, ultimately mapping the input to a 30×30 patch. It can simultaneously determine the authenticity of 30×30 image blocks in the image, and the average of the judgment results of all image blocks is used as the result of the image authenticity judgment.
[0050] Furthermore, the preprocessed underwater video dataset is input into the initial video enhancement model built based on CycleGAN for training. The optimized video enhancement model obtained includes:
[0051] S1, mapping the pre-processed underwater video data to generate underwater video enhancement data through a first generator;
[0052] S2. The underwater video enhancement data is discriminated by the discriminator. If the underwater video enhancement data is consistent with the pre-processed underwater video data, S3 is performed; if not, the discriminator parameters are modified and the return is returned to S1.
[0053] S3. Mapping the underwater video enhancement data to generate non-enhanced underwater video data through a second generator.
[0054] Furthermore, the loss function between the first generator and the discriminator is:
[0055]
[0056] Among them, X, Y represent the source domain and target domain respectively, p data (x) is the data distribution of source domain X, p data (y) is the data distribution of the target domain Y, E y ~p data (y) means y obeys p data (y) to find the expectation, Ex~p data (x) means x obeys p data (x) is the expected value, G represents the generator D Y Denotes the discriminator x as an element of the source domain X, y as an element of the target domain Y, and D Y (y) indicates whether the discriminator identifies y as belonging to the Y domain, G(x) indicates the input x element of the generator G, and D Y (G(x)) uses the discriminator to identify Y Don't input x into the G generator to see if the result of the input belongs to the Y domain.
[0057] Further, such as Figure 4 The cycle consistency loss between the first and second generators is shown to be:
[0058]
[0059] Among them, E y ~p data (y) means y obeys p data (y) to find the expectation, Ex~p data (x) means x obeys p data(x), G is a generator that converts elements of the X domain into the Y domain, F is a generator that converts elements of the Y domain into the X domain, x is an element of the X domain, y is an element of the Y domain, G(x) means using generator G to convert x into an image in the Y domain, F(G(x)) means first using generator G to convert x into an image in the Y domain, and then using generator F to convert the result of generator G back to the X domain, F(y) means using generator F to convert y into an image in the X domain, G(F(y)) means first using generator F to convert y into an image in the X domain, and then using generator G to convert the result of generator F back to the y domain.
[0060] Furthermore, the identity loss between the first generator and the second generator is:
[0061]
[0062] Among them, E y ~p data (y) means y obeys p data (y) to find the expectation, Ex~p data (x) means x obeys p data (x), G is a generator that converts elements of the X domain into the Y domain, F is a generator that converts elements of the Y domain into the X domain, x is an element of the X domain, y is an element of the Y domain, G(y) represents the image of y converted to the Y domain using the generator G, and F(x) represents the image of x converted to the X domain using the generator F.
[0063] An electronic device includes: a memory and a processor; the memory is used to store a program; the processor is used to execute the program to implement an underwater video enhancement method based on CycleGAN.
[0064] A readable storage medium stores a computer program, which, when executed by a processor, implements a CycleGAN-based underwater video enhancement method.
[0065] The experimental data used in this paper is the large-scale, high-resolution Underwater Video Enhancement Benchmark (UVEB), which was released on July 7, 2024. It contains 1,308 video sequences. UVEB comes from multiple countries and covers a variety of scenes and video degradation types to adapt to diverse and complex underwater environments.
[0066] Due to the large size of the UVEB dataset, videos range from 1 to 20 seconds in duration, with frame rates ranging from 30 to 60 frames, and resolutions ranging from 38% ultra-high-definition (UHD) to 62% 4K. In this paper, only the first 150 pairs of video sequences were used, consisting of 150 source and 150 target videos. Images were captured every 6 frames, resulting in a total of 16,836 images, including 8,418 source and 8,418 target sequences. The training set to test set ratio was 9:1, with 7,563 training images and 855 test images.
[0067] In this experiment, the truncated data was randomly cropped to 512×512 as the initial input to the model, which increased the training speed. The data was randomly shuffled and then randomly reversed before input. This increased the randomness and diversity of the learning data, improving the model's generalization ability.
[0068] like Figure 5 The figure shows the experimental results of different methods on real underwater video enhancement. Figure a shows RAW, b shows CycleGAN (old), c shows FA+Net, d shows UVE_Net, e shows the results of the present invention, and f shows GT. Three consecutive frames were extracted from the test video and are shown from left to right. The models tested included the original CycleGAN model. Due to its simplicity, the generator fell into a local optimum to deceive the discriminator, resulting in severe color deviation.
[0069] The indicators used in this experiment are: Peak Signal-to-Noise Ratio (PSNR), Mean Squared Difference Between Predicted and True Values (MSE), and Structural Similarity Index (SSIM). The calculation methods are as follows:
[0070] ,
[0071] ,
[0072] ,
[0073] Among them, MAXI represents the maximum value of the image color; MSE represents the mean square error; m represents the height of two images with the same height and width; n represents the width of two images with the same height and width; I and K refer to two images with the same height and width; j represents the horizontal and vertical coordinates of the pixel point in the image; I(i, j)-K(i, j) represents the difference in the pixel values at the same position in the two images; SSIM mainly considers three key features of the image: luminance, contrast, and structure; x and y are the two images to be compared; l(x, y) represents the calculation of luminance similarity; c(x, y) represents the calculation of contrast similarity; s(x, y) represents the calculation of structure similarity; Indicates the proportion of different features in the SSIM measurement.
[0074] The effects of each model in the model evaluation are as follows:
[0075] Methods PSNR↑ MSE↓ SSIM↑ CycleGAN(old) 27.87 106.17 0.066 FA+Net 28.33 95.39 0.95 UVE_Net 27.97 103.62 0.93 The present invention 29.01 81.55 0.951
[0076] Experiments show that the underwater video enhancement method based on CycleGAN proposed in this paper has a significant effect on improving the image generation quality.
[0077] The present invention discloses a method for underwater video enhancement based on CycleGAN. CycleGAN is used as a model for unsupervised training without paired data, which reduces the model's requirements for the data set while avoiding the mode collapse problem that is prone to CycleGAN. The present invention uses the UEVB large-scale high-resolution data set, has more excellent video enhancement capabilities, and can meet high-resolution input. The network structure in the present invention uses information interaction between frames, uses intermediate frame features as convolution kernels, and passes the information of the auxiliary network to the main network information, which speeds up the training speed while improving the accuracy of the model. The present invention has strong universality. Thanks to the powerful fitting ability of the network, even in the face of underwater situations under different circumstances, good results can be achieved after perfect training. The present invention implements the UVE_Net network as the generator of the CycleGAN model, changes the single loss function of the UVE_Net network, and reduces the requirements for the data set, thereby improving the training speed while improving the accuracy of the model and avoiding the mode collapse that is prone to CycleGAN models.
[0078] The above are merely preferred embodiments of the present application, but the scope of protection of the present application is not limited thereto. Any changes or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in this application should be included in the scope of protection of the present application. Therefore, the scope of protection of the present application should be based on the scope of protection of the claims.
Claims
1. A method for underwater video enhancement based on CycleGAN, characterized in that: include: Acquire an underwater video data set, and preprocess the underwater video data in the underwater video data set to obtain preprocessed underwater video data; The preprocessed underwater video dataset is input into the initial video enhancement model built based on CycleGAN for training to obtain the optimized video enhancement model; The initial video enhancement model constructed based on CycleGAN includes a generator network and a discriminator network, wherein the generator network includes a first generator and a second generator; After the original video frame sequence is input, the intermediate frame is first extracted and input into the first generator. The preliminary feature extraction is completed after four-fold downsampling and convolution. The intermediate extracted features are generated through 30 residual blocks. The first convolution kernel is generated through downsampling. At the same time, the first residual result is downsampled through 30 residual blocks to generate the second convolution kernel. In the second generator, the input frame sequence is first downsampled four times, and then subjected to three-dimensional convolution to extract features. After that, convolution is performed separately for each individual frame, and the convolution kernel is the first convolution kernel generated by the intermediate frame in the first generator; the first convolution result is upsampled twice and then passed through 15 three-dimensional residual blocks, and the second output generated in the first generator is used as the convolution kernel of each individual frame for convolution again to complete the information exchange between frames; In the discriminator network, the input frame sequence is divided into single frames for separate judgment by the discriminator, and the results of all frames are averaged as the final judgment result of the frame sequence. After the image sequence is input, the image sequence is divided into pictures by frame, and firstly, a preliminary feature extraction is performed through an initial convolution and activation layer, and then the convolution block is input. The structure of the convolution block is to perform a convolution on the input and then perform a one-fold downsampling, and then input the activation layer after instance normalization. Three convolution blocks are used continuously in the discriminator network, and each convolution block contains a downsampling, and finally three downsampling operations are realized; the output result is convolved and activated using a sigmoid function, and all frames are averaged after processing, and the result is divided into two to complete the judgment; The target underwater video data is obtained, and the optimized video enhancement model is input to obtain the target underwater video enhancement data.
2. The underwater video enhancement method based on CycleGAN according to claim 1, characterized in that: The preprocessed underwater video dataset is input into the initial video enhancement model built based on CycleGAN for training. The optimized video enhancement model includes: S1, mapping the pre-processed underwater video data to generate underwater video enhancement data through the first generator; S2, discriminating the underwater video enhancement data through the discriminator, if the underwater video enhancement data is consistent with the pre-processed underwater video data, proceed to S3; if they are inconsistent, modifying the discriminator parameters and returning to S1; S3. Mapping the underwater video enhancement data to generate unenhanced underwater video data through the second generator.
3. The underwater video enhancement method based on CycleGAN according to claim 1, characterized in that: The loss function between the first generator and the discriminator is: Among them, X, Y represent the source domain and target domain respectively, p data (x) is the data distribution of source domain X, p data (y) is the data distribution of the target domain Y, E y ~p data (y) means y obeys p data (y) to find the expectation, Ex~p data (x) means x obeys p data (x), G represents the generator, D Y represents the discriminator, x is an element of the source domain X, y is an element of the target domain Y, and D Y (y) indicates whether the discriminator identifies whether y belongs to the Y domain, G(x) indicates the input x element of the generator G, and D Y (G(x)) represents the discriminator D Y Don't input x into the G generator to see if the result of the input belongs to the Y domain.
4. The underwater video enhancement method based on CycleGAN according to claim 1, characterized in that: The cycle consistency loss between the first generator and the second generator is: Among them, E y ~p data (y) means y obeys p data (y) to find the expectation, Ex~p data (x) means x obeys p data (x), G is a generator that converts elements in the X domain to the Y domain, F is a generator that converts elements in the Y domain to the X domain, x is an element in the X domain, y is an element in the Y domain, G(x) means using generator G to convert x to an image in the Y domain, F(G(x)) means first using generator G to convert x to an image in the Y domain, and then using generator F to convert the result of generator G back to the X domain, F(y) means using generator F to convert y to an image in the X domain, G(F(y)) means first using generator F to convert y to an image in the X domain, and then using generator G to convert the result of generator F back to the y domain.
5. The underwater video enhancement method based on CycleGAN according to claim 1, characterized in that: The identity loss between the first generator and the second generator is: Among them, E y ~p data (y) means y obeys p data (y) to find the expectation, Ex~p data (x) means x obeys p data (x), G is a generator that converts elements in the X domain to the Y domain, F is a generator that converts elements in the Y domain to the X domain, x is an element in the X domain, y is an element in the Y domain, G(y) represents the image converted from y to the Y domain using the generator G, and F(x) represents the image converted from x to the X domain using the generator F.
6. An electronic device, characterized in that: include: Memory and processor; The memory is used to store programs; The processor is used to execute the program to implement the underwater video enhancement method based on CycleGAN as described in any one of claims 1-5.
7. A readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the underwater video enhancement method based on CycleGAN as described in any one of claims 1 to 5 is implemented.
Citation Information
Patent Citations
Underwater video image real-time enhancement method
CN112102186A
Human body abnormal behavior recognition method and device
CN113449679A