Underwater fish body depth image denoising restoration method
By using AOT-GAN model and ZED binocular camera technology to process the depth images of underwater fish bodies, the problem of difficulty in dealing with underwater depth images in the prior art is solved, efficient and accurate denoising and repairing of fish body depth images is achieved, and the performance of computer vision tasks is improved.
Patent Information
- Application Number
- CN202510031357.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-09
- Publication Date
- 2025-05-13
AI Technical Summary
The prior art is difficult to effectively process underwater depth images, especially dynamic fish body depth images, and there is a lack of a generative adversarial network model suitable for fish body depth images.
Generative adversarial network (GAN) model, especially the AOT-GAN model, combined with ZED binocular camera technology, accurately left and right views, calculate depth images, and efficiently produce data sets through the dataset building module to realize denoising and repairing fish body depth images.
It improves the efficiency and accuracy of fish body depth image processing, realizes the effective application of generative adversarial networks in the field of fish body depth maps, and significantly improves the performance of computer vision tasks.
Smart Images

Figure CN119991476A_ABST
Abstract
Description
Technical Field
[0001] The invention belongs to the technical field of underwater depth map image processing, and in particular relates to an underwater fish body depth image denoising and restoration method. Background Art
[0002] In the field of digital image processing technology, the methods for denoising and repairing images are mainly divided into two categories: one is the traditional denoising algorithm based on filtering, which has a high operating efficiency, and the other is the image restoration algorithm based on deep learning, which has a better restoration effect.
[0003] There are few studies on denoising and restoration of depth images, and most of them are processed using convolutional neural networks. Currently, there is no research using generative adversarial networks in the field of depth image denoising and restoration. Due to environmental differences, existing technologies cannot specifically process underwater depth maps, especially dynamic fish depth maps. In the prior art, the algorithm based on generative adversarial networks has a good restoration effect, but there is no model suitable for fish depth images in this field; and the construction of a generative adversarial network dataset is extremely difficult, and there is currently no public fish depth image dataset for training the model. For this reason, the present invention proposes a method for denoising and restoring underwater fish depth images. The present invention innovatively uses a generative adversarial network to complete depth information and specifically processes fish depth images. Summary of the invention
[0004] The purpose of the present invention is to provide a method for denoising and repairing underwater fish depth images, with the aim of producing a depth image dataset of underwater fish bodies, and using a generative adversarial network to denoise and repair the depth map, thereby creating an integrated model of dataset production and depth image restoration, thereby realizing the use of a generative adversarial network in the field of fish depth maps, and improving the efficiency and accuracy of computer vision tasks that require fish depth map information.
[0005] The technical solution adopted by the present invention is as follows:
[0006] A method for denoising and restoring an underwater fish depth image comprises the following steps:
[0007] Step 1: Complete data acquisition by importing the collected underwater fish video files;
[0008] Step 2: Configure the environmental parameters of the computer system for digital image processing and directly run the AOT-GAN model;
[0009] Step 3: Modify the pre-trained model and build the dataset building module;
[0010] Step 4: Prepare the data set and divide it into training set and test set;
[0011] Step 5: Use the training set and test set to train the AOT-GAN model;
[0012] Step 6: Improve and adjust the model according to the training results, and repeat step 5 until the evaluation indicators are met, and then output the final AOT-GAN model;
[0013] Step 7: Directly use the trained AOT-GAN model to denoise and repair the depth image of the underwater fish collected in real time.
[0014] Preferably, in step 1, an underwater drone equipped with a 3D sensing camera is placed at the bottom of a fish pond, and the underwater lighting of the drone is turned on to obtain a video of a fish body swimming freely in the water.
[0015] Preferably, in step 2, during the environment configuration process, the AOT-GAN model is directly run on a computer system with the environment configured.
[0016] Preferably, in step 3, the pre-trained model is modified to G0000000.pt, and the pconv partial convolution technology is used as the underlying data processing type of the model. Partial feature extraction is performed according to the mask label in the pconv partial convolution technology, and compared with the global features; the original video data is directly imported as the data input method, and a data set construction module is constructed in the model. The data set construction module first extracts frames from the video, batch saves and captures the effective depth information of the center part of the image, and modifies the image size to meet the standard of the model and saves it to a local folder; the data set construction module is also used to extract features from the original data set, batch produce label data sets that meet the standards of the pconv partial convolution technology, save them to a local folder and directly import them into the model training module.
[0017] Preferably, the step 4 specifically includes the following steps:
[0018] Step 401: firstly, the original fish swimming video is subjected to frame extraction, and the frames are extracted at a frequency of every 10 frames, and the frames are saved to a local folder;
[0019] Step 402: Then, the initial folder is screened for invalid data, and data that cannot be processed due to incomplete fish bodies and hardware reasons are screened out to form an initial data set;
[0020] Step 403: Process the initial data set, first extract the effective depth information of the central part, then modify the image size according to the model bottom layer data processing standard, and after processing, rename the batch and save it to a local folder;
[0021] Step 401: Create a mask data set, extract global features of the image, identify the missing depth information parts and create labels, and save the label data sets in batches;
[0022] Step 405: Use Python language to write a program to divide the image data set and labels, and randomly divide the training set and the test set into a ratio of 8:2.
[0023] Preferably, in step 5, the AOT-GAN model is used, batch-size is set to 2, mask_type is set to pconv partial convolution, image size is set to 512*512, and pre-trained model is set to G0000000.pt; if the data set has been created, the model is entered from the training module, the data set path has been passed in as a parameter setting, and the data set is directly imported into the module without creating a folder or modifying the path. In the data set creation module, only the video path needs to be passed in, and data set construction and model training can be integrated; during the training process, AOT-block is used for feature extraction, and the label data set combined with the pconv partial convolution technology is feature matched with the original data set. This method improves the efficiency and accuracy of model learning and improves the model's processing ability for fish depth images.
[0024] Preferably, in step 6, in the process of improving and adjusting the model according to the training results, for the fish body depth image, compared with the building data set originally applied to the model, the feature complexity is lower, and different modification schemes are tried to select the best related adjustment. The AOT-GAN network model generator neural network is set to 3 layers and the corresponding parameters are lowered, the stride convolution step size is set to 2, and the padding is set to 1. The model training speed and model operation efficiency are improved without affecting the processing accuracy; the improved model is retrained, and the training steps are the same as above.
[0025] Preferably, the evaluation index in step 6 is determined by a peak signal-to-noise ratio (PSNR), which is calculated based on a mean square error (MSE). The mean square error (MSE) measures the difference in pixel values between the original image and the processed image. The PSNR is a nonlinear transformation of the MSE, expressed in decibels. A higher value indicates a smaller distortion and a better image quality. The formula is:
[0026]
[0027] The root mean square error RMSE represents the square sum of the squares of the pixel difference between the result and the true value, and its formula is:
[0028]
[0029] In the formula, D i Represents the predicted result, Indicates the true value of the pixel;
[0030] The mean absolute error MAE represents the average value of the absolute error between the predicted depth image and the true value, and its formula is:
[0031]
[0032] The technical effects achieved by the present invention are:
[0033] The purpose of the present invention is to produce a depth image dataset of underwater fish bodies, and use a generative adversarial network to denoise and repair the depth map, create an integrated model of dataset production and depth image restoration, thereby realizing the use of a generative adversarial network in the field of fish depth maps, and improving the efficiency and accuracy of computer vision tasks that require fish depth map information. BRIEF DESCRIPTION OF THE DRAWINGS
[0034] Figure 1 It is a technical scheme diagram of a method for denoising and repairing an underwater fish body depth image of the present invention;
[0035] Figure 2 The present invention is an algorithm flow chart of an underwater fish body depth image denoising and restoration method. DETAILED DESCRIPTION
[0036] In order to make the purpose and advantages of the present invention more clearly understood, the present invention is specifically described below in conjunction with embodiments. It should be understood that the following text is only used to describe one or several specific embodiments of the present invention, and does not strictly limit the scope of protection of the specific claims of the present invention.
[0037] like Figure 1 as well as Figure 2 As shown, a method for denoising and repairing an underwater fish depth image comprises the following steps:
[0038] Step 1: Complete data acquisition by importing the collected underwater fish video files;
[0039] In step 1, an underwater drone equipped with a 3D sensing camera is placed at the bottom of a fish pond, and the underwater lighting of the drone is turned on to obtain a video of a fish body swimming freely in the water;
[0040] In this embodiment, the video data of the spotted sea bream is:
[0041] The video collection was completed in Laizhou Mingbo Aquatic Products. The collected images came from a fish pond with a diameter of 5 meters and a depth of 80 centimeters. The water in the fish pond came from filtered seawater. An underwater drone equipped with ZED's second-generation 3D sensing camera was placed at the bottom of the pond. The underwater lighting on the drone was turned on to obtain images of fish swimming freely in the water. The camera was called through ZED's SDK to initialize the camera. mm was used as the length unit, the distance range was set to 200mm to 20000mm, the camera's built-in image enhancement was turned on, the depth stabilization algorithm was turned on, the camera depth confidence was set to 100, the fps was set to 30, the resolution was 1920*1080*2, the video format was svo, and the video was recorded for 10 hours.
[0042] In this embodiment, largemouth bass video data:
[0043] The video collection was completed at the National Digital Fisheries Innovation Center. The collected images came from a fish pond with a diameter of 3 meters and a depth of 45 centimeters. An underwater drone equipped with ZED's second-generation 3D sensing camera was placed at the bottom of the pond to obtain images of fish swimming freely in the water. The camera was called through ZED's SDK to initialize the camera, using mm as the length unit, the distance range was set to 200mm to 20000mm, the camera's built-in image enhancement was turned on, the depth stabilization algorithm was turned on, the camera depth confidence was set to 100, the fps was set to 30, the resolution was 1920*1080*2, the video format was svo, and the video was recorded for 5 hours.
[0044] Step 2: Configure the environmental parameters of the computer system for digital image processing and directly run the AOT-GAN model;
[0045] In step 2, during the environment configuration process, the AOT-GAN model is directly run on the computer system with the environment configuration;
[0046] In this embodiment, the AOT-GAN model runs in the local environment, the local environment is the win11 operating system, the program version is written in python 3.8.8, the memory capacity is 16GB, the hard disk capacity is 1TB, the torch version is 1.8.0, the CUDA version is 10.2, the imageio version is 2.9.0, the matplotlib version is 3.4.1, the opencv-python language version is 4.5.1, and the ZED SDK version is 3.7.
[0047] Step 3: Modify the pre-trained model and build the dataset building module;
[0048] In the step 3, the pre-trained model is modified to G0000000.pt, and the pconv partial convolution technology is used as the underlying data processing type of the model. Partial feature extraction is performed according to the mask label in the pconv partial convolution technology, and compared with the global features; the original video data is directly imported as the data input method, and a data set construction module is constructed in the model. The data set construction module first extracts frames from the video, saves and captures the effective depth information of the center part of the image in batches, and modifies the image size to meet the model standards and saves it to a local folder; the data set construction module is also used to extract features from the original data set, batch produce label data sets that meet the pconv partial convolution technology standards, save them to a local folder and directly import them into the model training module.
[0049] Step 4: Prepare the data set and divide it into training set and test set;
[0050] Step 401: firstly, the original fish swimming video is subjected to frame extraction, and the frames are extracted at a frequency of every 10 frames, and the frames are saved to a local folder;
[0051] Step 402: Then, the initial folder is screened for invalid data, and data that cannot be processed due to incomplete fish bodies and hardware reasons are screened out to form an initial data set;
[0052] Step 403: Process the initial data set, first extract the effective depth information of the central part, then modify the image size according to the model bottom layer data processing standard, and after processing, rename the batch and save it to a local folder;
[0053] Step 401: Create a mask data set, extract global features of the image, identify the missing depth information parts and create labels, and save the label data sets in batches. In the present invention, Mask is the name of one of the data sets;
[0054] Step 405: Use Python language to write a program to divide the image data set and labels, and randomly divide the training set and the test set into a ratio of 8:2.
[0055] Step 5: Use the training set and test set to train the AOT-GAN model;
[0056] In the step 5, the AOT-GAN model is used, batch-size is set to 2, mask_type is set to pconv partial convolution, image size is set to 512*512, and the pre-trained model is set to G0000000.pt; if the data set has been created, the model is entered from the training module, the data set path has been passed in as a parameter setting, and the data set is directly imported into the module without creating a folder or modifying the path. In the data set creation module, only the video path needs to be passed in, and data set construction and model training can be integrated; during the training process, AOT-block is used for feature extraction, and the label data set combined with the pconv partial convolution technology is feature matched with the original data set. This method improves the efficiency and accuracy of model learning and improves the model's processing ability for fish depth images.
[0057] In the present invention, "batch-size" is a parameter in the model, which means: "the number of samples selected for one training"; "mask_type" is another parameter in the model, which means: "label type".
[0058] Step 6: Improve and adjust the model according to the training results, and repeat step 5 until the evaluation indicators are met, and then output the final AOT-GAN model;
[0059] In step 6, in the process of improving and adjusting the model according to the training results, for the fish body depth image, compared with the building data set originally applied to the model, the feature complexity is lower, and different modification schemes are tried to select the best related adjustment. The AOT-GAN network model generator neural network is set to 3 layers and the corresponding parameters are lowered, the stride convolution step size is set to 2, and the padding is set to 1. The model training speed and model operation efficiency are improved without affecting the processing accuracy; the improved model is retrained, and the training steps are the same as above.
[0060] The evaluation index described in step 6 is determined by the peak signal-to-noise ratio (PSNR), which is calculated based on the mean square error (MSE). The mean square error (MSE) measures the difference in pixel values between the original image and the processed image. The PSNR is a nonlinear transformation of the MSE, expressed in decibels. The higher the value, the smaller the distortion and the better the image quality. The formula is:
[0061]
[0062] The root mean square error RMSE represents the square sum of the squares of the pixel difference between the result and the true value, and its formula is:
[0063]
[0064] In the formula, D i Represents the predicted result, Indicates the true value of the pixel;
[0065] The mean absolute error MAE represents the average value of the absolute error between the predicted depth image and the true value, and its formula is:
[0066]
[0067] Step 7: Directly use the trained AOT-GAN model to denoise and repair the depth images of underwater fish collected in real time.
[0068] The core of the present invention is to innovatively propose a method that integrates the production of fish body depth image datasets and image denoising and restoration. This method is rooted in the advanced concept of generative adversarial networks, especially the use of the AOT-GAN model and the integration of ZED binocular camera technology to obtain accurate left and right views, and then calculate the depth image. By optimizing and adjusting the AOT-GAN model in a targeted manner, it is adapted to the processing requirements of fish body depth images. On this basis, the present invention cleverly integrates a batch dataset acquisition module, thereby realizing an integrated process of efficient dataset production and depth image denoising and restoration.
[0069] In order to verify and improve the performance of the model, I used the datasets of spotted sea bream and largemouth bass to fully train the model, and carefully fine-tuned the model based on the training feedback. This series of efforts enabled the model to excellently complete the task of completing the depth information in the depth image of the fish body.
[0070] Compared with pure computer theoretical algorithms, the present invention is tailored for specific application scenarios. It can accurately and efficiently process the depth image of fish in factory farming environments, showing a high degree of pertinence and practicality. The present invention has a wide range of applications, especially in computer vision tasks that rely on fish depth maps, which can significantly improve the efficiency and accuracy of task execution.
[0071] The present invention proposes a method based on generative adversarial networks, which shows more outstanding effects in denoising and repairing fish depth images, and can achieve accurate processing of fish depth images. At the same time, the present invention also innovatively adds a dataset construction module to the model. This design effectively alleviates the problem of restricting the application of generative adversarial networks due to the difficulty of dataset construction in this field, and provides a new solution for image processing in related fields.
[0072] In the present invention, the Chinese name of “AOT-GAN model” is: “Contextual Information Aggregation Transformation Generative Adversarial Network”;
[0073] The above is only a preferred embodiment of the present invention. It should be noted that, for those skilled in the art, several improvements and modifications can be made without departing from the principles of the present invention, and these improvements and modifications should also be considered as the protection scope of the present invention. The structures, devices and operating methods not specifically described and explained in the present invention shall be implemented according to the conventional means in the art unless otherwise specified and limited.
Claims
1. A method for denoising and restoring underwater fish depth images, characterized in that: The following steps are involved: Step 1: Complete data acquisition by importing the collected underwater fish video files; Step 2: Configure the environmental parameters of the computer system for digital image processing and directly run the AOT-GAN model; Step 3: Modify the pre-trained model and build the dataset building module; Step 4: Prepare the data set and divide it into training set and test set; Step 5: Use the training set and test set to train the AOT-GAN model; Step 6: Improve and adjust the model according to the training results, and repeat step 5 until the evaluation indicators are met, and then output the final AOT-GAN model; Step 7: Directly use the trained AOT-GAN model to denoise and repair the depth images of underwater fish collected in real time.
2. The method for denoising and restoring an underwater fish depth image according to claim 1, characterized in that: In the step 1, an underwater drone equipped with a 3D sensing camera is placed at the bottom of a fish pond, and the underwater lighting of the drone is turned on to obtain a video of a fish body swimming freely in the water.
3. The method for denoising and restoring an underwater fish depth image according to claim 1, characterized in that: In step 2, during the environment configuration process, the AOT-GAN model is directly run on the computer system with the environment configured.
4. The method for denoising and restoring an underwater fish depth image according to claim 1, characterized in that: In the step 3, the pre-trained model is modified to G0000000.pt, and the pconv partial convolution technology is used as the underlying data processing type of the model. Partial feature extraction is performed according to the mask label in the pconv partial convolution technology, and compared with the global features; the original video data is directly imported as the data input method, and a data set construction module is constructed in the model.
5. The method for denoising and restoring an underwater fish depth image according to claim 1, characterized in that: The step 4 specifically includes the following steps: Step 401: firstly, extract frames from the original fish swimming video and save it to a local folder; Step 402: Then, the initial folder is screened for invalid data, and data that cannot be processed due to incomplete fish bodies and hardware reasons are screened out to form an initial data set; Step 403: Process the initial data set, first extract the effective depth information of the central part, then modify the image size according to the model bottom layer data processing standard, and after processing, rename the batch and save it to a local folder; Step 401: Create a mask data set, extract global features of the image, identify the missing depth information parts and create labels, and save the label data sets in batches; Step 405: Use Python language to write a program to divide the image data set and labels, and randomly divide the training set and the test set into a ratio of 8:
2.
6. The method for denoising and restoring an underwater fish depth image according to claim 1, characterized in that: In step 5, the AOT-GAN model is used, batch-size is set to 2, mask_type is set to pconv partial convolution, image size is set to 512*512, and the pre-trained model is set to G0000000.pt; if the data set has been created, the model is entered from the training module, the data set path has been passed in as a parameter setting, and the data set is directly imported into the module. During the training process, the AOT-block model is used for feature extraction, and the label data set combined with the pconv partial convolution technology is feature matched with the original data set.
7. The method for denoising and restoring an underwater fish depth image according to claim 1, characterized in that: In step 6, in the process of improving and adjusting the model according to the training results, the AOT-GAN network model generator neural network is set to 3 layers and the corresponding parameters are lowered, the stride convolution step size is set to 2, the padding is set to 1, and the improved model is retrained, and the training steps are the same as above.
8. The method for denoising and restoring an underwater fish depth image according to claim 7, characterized in that: The evaluation index described in step 6 is determined by the peak signal-to-noise ratio PSNR, and the peak signal-to-noise ratio PSNR is calculated based on the mean square error MSE; The formula is: The root mean square error RMSE represents the square sum of the squares of the pixel difference between the result and the true value, and its formula is: In the formula, D i Represents the predicted result, Indicates the true value of the pixel; The mean absolute error MAE represents the average value of the absolute error between the predicted depth image and the true value, and its formula is: The peak signal-to-noise ratio (PSNR) is calculated based on the mean square error (MSE). The formula is: Where MAX is the maximum pixel value of the image, MSE is the mean square error between the original image and the reconstructed image; m and n are the number of rows and columns of the image, respectively; I(i, j) is the pixel value of the original image at position (i, j); K(i, j) is the pixel value of the comparison image at position (i, j).