Underwater netting operation robot image compression and decompression method and device

By introducing frequency domain cross-scale fusion module and deredundant module into the underwater image compression network, the problems of low underwater image compression efficiency and poor reconstruction quality in the prior art are solved, and efficient and low bit rate underwater image compression and reconstruction are achieved.

CN120050425AActive Publication Date: 2025-05-27HUNAN UNIV
View PDF 13 Cites 0 Cited by

Patent Information

Application Number
CN202510531138.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-25
Publication Date
2025-05-27
Estimated Expiration
2045-04-25

AI Technical Summary

Technical Problem

The existing image compression network has low compression efficiency and poor reconstruction quality in underwater image processing, making it difficult to take into account the global correlation between features, affecting the development of subsequent tasks.

Method used

The underwater image compression model based on frequency domain cross-scale fusion is adopted, and the compression performance and reconstruction quality of the model are improved through the scale hyper-priority architecture and the frequency domain cross-scale fusion module, combined with the de-redundant module.

Benefits of technology

It improves the compression efficiency and reconstruction quality of underwater images, realizes low bit rate and high-fidelity compression in different environments, reduces the pressure of underwater transmission, and supports the development of underwater mesh operation robot tasks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120050425A_ABST
    Figure CN120050425A_ABST
Patent Text Reader

Abstract

The invention discloses an underwater netting operation robot image compression and decompression method and device, and the method comprises the steps: firstly constructing an underwater image compression model based on frequency domain cross-scale fusion, which comprises a compression module and a decompression module, and specifically comprises two encoders, two quantization modules, two arithmetic encoders, two arithmetic decoders and two decoders; wherein the first encoder comprises a frequency domain cross-scale fusion module, and each of the first encoder and the first decoder comprises a redundancy elimination module; then training the constructed underwater image compression-decompression model by using an underwater netting image training set; using the trained compression module to compress the image acquired by the underwater netting operation robot to obtain corresponding bit stream data; and using the trained decompression module to reconstruct the underwater image from the bit stream data after the compression of the obtained underwater image. According to the method, the feature extraction capability is improved by capturing complementary information of different receptive fields, meanwhile, the influence of redundant information is reduced, and the image compression performance is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of image compression encoding, and particularly relates to a method and device for compressing and decompressing images of an underwater netting operation robot. Background Art

[0002] In the deep - sea cage aquaculture industry, using an underwater netting operation robot to replace traditional manual labor for netting cleaning has safety and high efficiency, which is the industry trend. The operator remotely controls the underwater netting operation robot to complete the cleaning work through a human - machine interaction visualization interface on the ground. Considering the limited underwater bandwidth, the underwater images collected by the underwater netting operation robot need to be compressed at a low bit rate and then transmitted to the ground receiving end for reconstruction. Considering subsequent cleaning tasks, the underwater netting images require a high reconstruction quality.

[0003] Underwater imaging has uniqueness and complexity. Due to the scattering of light and the influence of suspended particles, the clarity of underwater images is poor, and the disturbance of water flow and the movement of the underwater netting operation robot further reduce the detail quality of underwater images. Most existing image compression networks achieve feature extraction through multi - layer stacking of convolutional kernels with small receptive fields, making it difficult to take into account the global correlation between features, which leads to problems of low compression efficiency and poor reconstruction quality on underwater images, and is not conducive to the development of subsequent tasks. Summary of the Invention

[0004] The present invention provides a method and device for compressing and decompressing images of an underwater netting operation robot, which can improve the compression efficiency and reconstruction quality.

[0005] To achieve the above - mentioned technical objectives, the present invention adopts the following technical solutions: A method for compressing and decompressing images of an underwater netting operation robot, comprising: Step 1, constructing an underwater image compression model based on frequency - domain cross - scale fusion; The underwater image compression - decompression model adopts a scale hyper - prior architecture, and sequentially includes a first encoder, a first quantization module, a first arithmetic encoder, a first arithmetic decoder, and a first decoder from input to output; wherein the first encoder includes a frequency - domain cross - scale fusion module, and both the first encoder and the first decoder include a redundancy removal module; At the output end of the first encoder, a second encoder, a second quantization module, a second arithmetic encoder, a second arithmetic decoder, and a second decoder are further sequentially arranged; The first encoder, the first quantization module, the first arithmetic encoder, the second encoder, the second quantization module, the second arithmetic encoder, and the second decoder together constitute the compression module of the underwater image compression-decompression model; the first arithmetic decoder, the first decoder, the second arithmetic decoder, and the second decoder together constitute the decompression module of the underwater image compression-decompression model; Step 2: Use the underwater netting image training set to train the constructed underwater image compression-decompression model; Step 3: Use the trained compression module to compress the images collected by the underwater netting operation robot to obtain the corresponding bitstream data; use the trained decompression module to reconstruct the underwater image from the obtained bitstream data after compressing the underwater image.

[0006] Furthermore, in the compression module of the underwater image compression-decompression model: The first encoder is used to map the underwater image x collected by the underwater netting operation robot to a low-dimensional latent space for feature extraction and output the latent representation y; The second encoder is used to capture the spatial dependence between different elements in the latent representation y and output the low-dimensional latent variable ; The second quantization module is used to perform quantization operations on the continuous latent variable to obtain the quantized variable ; The second arithmetic encoder is used to encode the quantized variable according to its own probability distribution to obtain the second bitstream ; The first quantization module is used to perform quantization operations on the continuous latent representation y to obtain the quantized representation ; The second decoder is used to dynamically estimate the standard deviation of the quantized representation according to the quantized variable , construct a Gaussian probability model with a mean of 0 and a standard deviation of ; The first arithmetic encoder is used to encode the quantized representation according to the Gaussian probability model to obtain the first bitstream ; where the first bitstream and the second bitstream together constitute the bitstream data after compressing the underwater image; ; and / or, In the decompression module of the underwater image compression-decompression model: The second arithmetic decoder is used to decode the quantized variable according to the obtained second bitstream ; ; The second decoder dynamically estimates the standard deviation of the quantization representation according to the quantization variable and constructs a Gaussian probability model with a mean of 0 and a standard deviation of ; The first arithmetic decoder decodes the obtained first bitstream into a quantization representation according to the Gaussian probability model ; The first decoder then reconstructs the quantization representation to obtain the underwater image. ; Furthermore, the underwater netting image training set is sourced from real deep - sea cage farming scenarios, and underwater netting images at different depths, environments, and resolutions are selected.

[0007] Furthermore, the frequency - domain cross - scale fusion module is implemented by setting a spatial - domain path and a frequency - domain path and through information complementarity between the two paths:

[0008] First, the input feature map F is divided into a local feature map and a global feature map along the feature channel dimension according to a ratio ; Then, in the spatial - domain path, the local feature map and the global feature map are respectively convolution - processed and then fused to obtain a local output feature map ; In addition, in the frequency - domain path, on the one hand, the local feature map is convolution - processed, and on the other hand, the global feature map is downsampled, convolution - processed, and Fourier - transformed to the frequency domain for processing. Then, the feature maps obtained from both aspects are fused to obtain a global output feature map ; Finally, the local output feature map and the global output feature map are concatenated along the feature channel dimension to obtain the output feature map M of the frequency - domain cross - scale fusion module.

[0009] Furthermore, the steps of the spatial - domain path are as follows: The local feature map is convolved through a normal convolution with an input channel of and an output channel of to obtain a feature map , and the global feature map is convolved through a normal convolution with an input channel of and an output channel of to obtain a feature map ​​, the local output feature map can be obtained : ; Among them, represents the number of channels of the feature map F before the input frequency-domain cross-scale fusion module, represents the number of channels of the feature map M output from the frequency-domain cross-scale fusion module, and are both hyperparameters, is the proportion of local features in the input channels, is the proportion of local outputs in the output channels.

[0010] Furthermore, the local feature map is passed through a common convolution with an input channel of and an output channel of to obtain the feature map ; The global feature map is first downsampled through an average pooling layer, and after reducing the height and width by half, it passes through a convolutional block including a convolutional layer, a BN layer, and a ReLU activation function, and then through three branches respectively: the first branch first passes the feature map through the fast Fourier transform, extracts the real and imaginary parts of the complex result and concatenates them, then passes through a convolutional layer with a convolution kernel of 1×1, a BN layer, and a ReLU activation function, and finally separates the real and imaginary parts again and performs the inverse fast Fourier transform, retaining the real part as the output; the second branch outputs directly without any processing; the third branch first divides the feature map into 4 parts along half of the height and width respectively and reassembles them into a feature map along the channel dimension, processes them through the Fourier unit module and replicates and stacks them 4 times, and finally outputs a feature map with the same height and width as the input feature map of this branch; then the outputs of the three branches are added together and passed through a convolutional layer with a convolution kernel of 1×1 to obtain the feature map ; among them, the processing method of the first branch is denoted as the processing of the Fourier unit module; Finally, the feature maps and obtained from the frequency-domain path are added together to obtain the global output feature map : .

[0011] Furthermore, the redundancy removal module is applied after the last convolutional layer of the first encoder and before the first convolutional layer of the first decoder; for the redundancy removal module, first copy the input feature map, then initialize a 3×3 convolutional layer, apply this convolutional layer to some channels of the input feature map, and keep the remaining channels unchanged, and perform forward propagation in a sliced manner.

[0012] An underwater netting operation robot image compression device, which is composed of a compression module trained by the compression and decompression method described in any one of the above; The first encoder is used to map the underwater image x collected by the underwater netting operation robot to a low-dimensional latent space for feature extraction and output a latent representation y; The second encoder is used to capture the spatial dependence between different elements in the latent representation y and output a low-dimensional latent variable z; The second quantization module is used to perform quantization operations on the continuous latent variable z to obtain a quantized variable ; The second arithmetic encoder is used to encode the quantized variable according to its own probability distribution to obtain a second bitstream ; The first quantization module is used to perform quantization operations on the continuous latent representation y to obtain a quantized representation ; The second decoder is used to dynamically estimate the standard deviation of the quantized representation according to the quantized variable , construct a Gaussian probability model with a mean of 0 and a standard deviation of ; The first arithmetic encoder is used to encode the quantized representation according to the Gaussian probability model to obtain a first bitstream ; Among them, the first bitstream and the second bitstream together constitute the bitstream data after underwater image compression.

[0013] An underwater netting operation robot image decompression device, which is composed of a decompression module trained by the compression and decompression method described in any one of the above; The second arithmetic decoder is used to decode the quantized variable from the obtained second bitstream ; The second decoder dynamically estimates the standard deviation of the quantized representation according to the quantized variable , construct a Gaussian probability model with a mean of 0 and a standard deviation of ; The first arithmetic decoder decodes the obtained first bitstream into a quantized representation according to the Gaussian probability model; The first decoder then reconstructs the quantized representation to obtain an underwater image.

[0014] Compared with the prior art, the present invention has the following beneficial effects: (1) The computational complexity and the number of parameters of the frequency-domain cross-scale fusion module in the present invention are comparable to those of ordinary convolution. By the information complementarity between the spatial-domain path and the frequency-domain path, the receptive field is enlarged, and richer information is transmitted, enabling the model to obtain better compression performance.

[0015] (2) The present invention inserts a redundancy removal module into the first encoder and the first decoder, applies ordinary convolution to some input channels, and keeps the remaining channels unchanged, reducing redundant calculations and memory access, and further improving the compression ratio and reconstruction effect.

[0016] (3) The present invention has strong practicability, can realize low-bitrate and high-fidelity compression of underwater netting images in different environments, reduces the pressure of underwater transmission, and is conducive to the development of tasks of underwater netting operation robots and the subsequent detection work. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] Figure 1 It is a schematic diagram of the overall framework of an embodiment of the present invention.

[0018] Figure 2 It is a schematic diagram of the frequency-domain cross-scale fusion module of an embodiment of the present invention.

[0019] Figure 3 It is a schematic diagram of the compression process of an underwater netting operation robot according to an embodiment of the present invention.

[0020] Figure 4 It is a schematic diagram of the decompression process of a ground server according to an embodiment of the present invention.

[0021] Figure 5 It is a schematic diagram of the comparison of rate-distortion results according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0022] The embodiments of the present invention will be described in detail below. Based on the technical solution of the present invention, detailed implementation manners and specific operation processes are given, and the technical solution of the present invention is further explained and illustrated.

[0023] Embodiment 1

[0024] This embodiment provides a method for compressing and decompressing underwater netting operation robot images, including the following steps: Step 1, construct an underwater image compression model based on frequency-domain cross-scale fusion.

[0025] The underwater image compression - decompression model adopts a scale hyper - prior architecture and sequentially includes a first encoder, a first quantization module, a first arithmetic encoder, a first arithmetic decoder, and a first decoder from input to output. At the output end of the first encoder, a second encoder, a second quantization module, a second arithmetic encoder, a second arithmetic decoder, and a second decoder are also sequentially arranged. As Figure 1 shown.

[0026] 1. Compression module.

[0027] The first encoder, the first quantization module, the first arithmetic encoder, the second encoder, the second quantization module, the second arithmetic encoder, and the second decoder in the underwater image compression - decompression model together constitute the compression module.

[0028] The first encoder maps the underwater image x collected by the underwater netting operation robot to a low - dimensional latent space, extracts features and outputs a latent representation y.

[0029] The second encoder captures the spatial dependencies between different elements in the latent representation y and outputs a low - dimensional latent variable .

[0030] The second quantization module performs a quantization operation on the continuous latent variable to obtain a quantized variable .

[0031] The second arithmetic encoder encodes the quantized variable according to its own probability distribution to obtain a second bitstream .

[0032] The first quantization module performs a quantization operation on the continuous latent representation y to obtain a quantized representation .

[0033] The second decoder dynamically estimates the spatial distribution of the quantized representation according to the quantized variable to obtain the standard deviation of , and reconstructs a Gaussian probability model with a mean of 0 and a standard deviation of for each point in the latent representation, providing prior information for the first arithmetic encoder and the first arithmetic decoder.

[0034] The first arithmetic encoder encodes the quantized representation according to the Gaussian probability model to obtain a first bitstream , assigning short sequences to high probabilities and long sequences to low probabilities.

[0035] Among them, the first bitstream obtained by the first arithmetic encoder , together with the second bitstream obtained by the second arithmetic encoder , jointly constitute the bitstream data after underwater image compression.

[0036] 2. Decompression module.

[0037] The first arithmetic decoder, the first decoder, the second arithmetic decoder, and the second decoder in the underwater image compression-decompression model jointly constitute the decompression module.

[0038] The second arithmetic decoder decodes the quantization variable according to the obtained second bitstream ; ; The second decoder dynamically estimates the standard deviation of the quantization representation according to the quantization variable and constructs a Gaussian probability model with a mean of 0 and a standard deviation of ; ; The first arithmetic decoder decodes the obtained first bitstream into a quantization representation according to the Gaussian probability model ; ; The first decoder then reconstructs the quantization representation to obtain the underwater image.

[0039] 3. Frequency-domain cross-scale fusion module.

[0040] The first encoder of the present invention includes a frequency-domain cross-scale fusion module, that is, the structure of the first encoder includes a convolutional layer A, a non-linear normalization layer GDN, a frequency-domain cross-scale fusion module, GDN, a frequency-domain cross-scale fusion module, GDN, a convolutional layer B, and a redundancy removal module. The frequency-domain cross-scale fusion module is realized by setting a spatial-domain path and a frequency-domain path and through information complementarity between the two paths. As Figure 2 shown, the steps of the frequency-domain cross-scale fusion module are as follows: First, divide the input feature map F into a local feature map and a global feature map along the feature channel dimension according to the ratio : ; ; ; where H×W represents the spatial resolution of the input feature map, represents the number of channels of the input feature map, .

[0041] Then, in the spatial-domain path, the local feature map passes through an input channel of , the output channel is and a feature map is obtained through a common convolution with a convolution kernel of 5×5 , and the global feature map is passed through an input channel of , the output channel is and a feature map is obtained through a common convolution with a convolution kernel of 5×5 , and then the local output feature map can be obtained: ; Among them, represents the number of channels of the feature map F before the input frequency-domain cross-scale fusion module, represents the number of channels of the feature map M output from the frequency-domain cross-scale fusion module, and are both hyperparameters, is the proportion of local features in the input channel, is the proportion of local output in the output channel, and .

[0042] In addition, in the frequency-domain path, on the one hand, the local feature map is convolved, and on the other hand, the global feature map is downsampled, convolved, and Fourier-transformed to the frequency domain for processing, and then the feature maps obtained from both aspects are fused to obtain the global output feature map . The steps of the frequency-domain path are as follows: (1) Pass the local feature map through an input channel of , the output channel is and a feature map is obtained through a common convolution with a convolution kernel of 5×5 ; (2) First, downsample the global feature map through an average pooling layer, reduce the height and width by half, then pass through a convolution block including a convolution layer, a BN layer, and a ReLU activation function, and then pass through three branches respectively: the first branch first Fourier-transforms the feature map, extracts the real and imaginary parts of the complex result and splices them, then passes through a convolution layer with a convolution kernel of 1×1, a BN layer, and a ReLU activation function, and finally separates the real and imaginary parts again and performs an inverse fast Fourier transform, retaining the real part as the output; the second branch outputs directly without any processing; the third branch first divides the feature map into 4 parts along half of the height and width respectively and reassembles them into a feature map along the channel dimension, processes them through a Fourier unit module similar to the first branch, and then stacks and copies 4 parts for output, and finally outputs a feature map with the same height and width as the input feature map of this branch; finally, add the outputs of these three branches and pass through a convolution layer with a convolution kernel of 1×1 to obtain the feature map 。

[0043] (3) Add the feature maps obtained from the frequency-domain path and to obtain the global output feature map : 。

[0044] Finally, concatenate the local output feature map and the global output feature map along the feature channel dimension to obtain the output feature map M of the frequency-domain cross-scale fusion module.

[0045] 4. Redundancy removal module.

[0046] Both the first encoder and the first decoder of the present invention include a redundancy removal module, which is arranged after the last convolutional layer in the first encoder and before the first convolutional layer in the first decoder.

[0047] For the redundancy removal module, first copy the input feature map, then initialize a 3×3 convolutional layer, apply the convolutional layer to some channels of the input feature map, keep the remaining channels unchanged, and perform forward propagation in a sliced manner.

[0048] Step 2: Use the underwater netting image training set to train the constructed underwater image compression-decompression model.

[0049] The underwater netting image training set in this embodiment is sourced from the real deep-sea cage farming scenarios photographed by the BlueRov underwater netting operation robot at the Mindou No. 1 deep-sea farming platform. Underwater netting images at different depths, environments, and resolutions are selected, and a total of 2,273 images are collected. Among them, 2,000 images are used as the training set and 273 images are used as the test set, which are respectively used for the training and evaluation of the model.

[0050] The loss function for training the underwater image compression-decompression model in this embodiment aims to minimize the difference between the reconstructed image and the input image and minimize the bit rate required for encoding, so that the model can achieve the best performance under the rate-distortion constraint. The reconstruction error is the difference between the reconstructed image and the input image, which is calculated by the mean square error MSE. The rate loss is calculated by the information entropy, which includes the number of bits required for encoding the quantization representation and the number of bits required for encoding the quantization variable . The loss function is: ; where represents the loss function, is a hyperparameter used to balance the compression ratio and the image reconstruction quality. The smaller it is, the smaller the compressed bitstream file is, and the worse the image reconstruction effect obtained by decompression is; The larger it is, the larger the compressed bitstream file is, and the better the image reconstruction effect obtained by decompression is. E represents the reconstruction error, represents the compression code rate, , represents the number of bits required for encoding, represents the number of bits required for encoding.

[0051] In order for the gradient to be backpropagated during training, so as to optimize the parameters through gradient descent to make the model reach the best performance. The quantization module during the training process uses additive uniform noise to replace the quantizer to achieve approximate quantization operations: ; ; wherein, represents the mean noise with a value range of , represents the result of approximately quantizing the latent representation y, represents the result of approximately quantizing the latent variable z.

[0052] While the quantization module during testing and actual compression directly uses the quantizer for quantization operations: ; ;

[0053] wherein, represents the result of actual quantization of y, represents the result of actual quantization of z, represents quantization.

[0054] Step 3, use the trained compression module to compress the images collected by the underwater netting operation robot to obtain the corresponding compressed bitstream data; use the trained decompression module to reconstruct the underwater images from the obtained compressed bitstream data of the underwater images.

[0055] Embodiment 2

[0056] This embodiment provides a compression device for underwater netting operation robot images, which is applied to the underwater netting operation robot to compress the underwater images after collection to obtain a bitstream data file, facilitating subsequent storage or transmission. The compression device in this embodiment is composed of the trained compression module in the compression and decompression methods described in Embodiment 1, and includes: a first encoder, a first quantization module, a first arithmetic encoder, a second encoder, a second quantization module, a second arithmetic encoder, and a second decoder.

[0057] In this embodiment, the process of using the trained compression module to compress the images collected by the underwater netting operation robot is as follows Figure 3 as shown: The underwater image x collected by the underwater netting operation robot passes through the first encoder to map the underwater image x to a low-dimensional latent space for feature extraction and output a latent representation y; the second encoder captures the spatial dependence between different elements in the latent representation y and outputs a low-dimensional latent variable z; then, a quantizer is used to perform quantization operations on the continuous latent representation y and the continuous latent variable z respectively to obtain discrete and discrete ; the arithmetic encoder converts into a bitstream ; the quantized variable is used to dynamically estimate the standard deviation of through the second decoder, thereby establishing a probability model; the arithmetic encoder encodes the discrete into a bitstream according to its probability model; together constitute the bitstream data after underwater image compression for subsequent storage or transmission.

[0058] Embodiment 3

[0059] This embodiment provides a decompression device for underwater netting operation robot images, which is applied to a ground server to decode the received underwater image bitstream data into a reconstructed image and is composed of the decompression module trained in the compression and decompression method described in Embodiment 1, including: a first arithmetic decoder, a first decoder, a second arithmetic decoder, and a second decoder.

[0060] The second arithmetic decoder is used to decode the quantized variable from the obtained second bitstream ; The second decoder dynamically estimates the standard deviation of the quantized representation according to the quantized variable , constructs a Gaussian probability model with a mean of 0 and a standard deviation of ; The first arithmetic decoder decodes the obtained first bitstream into a quantized representation according to the Gaussian probability model; The first decoder then reconstructs the underwater image from the quantized representation .

[0061] The process of decompression is as follows Figure 4 as shown: The bitstream file is decoded by the arithmetic decoder to obtain discrete , the quantization variable is dynamically estimated by a second decoder for the standard deviation , thereby reconstructing a probability model, and an arithmetic decoder decodes a bitstream into a quantized representation , and the discrete is reconstructed into an input image by a first decoder.

[0062] The effects of the present invention are tested, and the results are as Figure 5 shown. The vertical axis is the peak signal-to-noise ratio PSNR, and the horizontal axis is the bit rate bpp. This graph is compared with the compression performance of JPEG2000. The higher the peak signal-to-noise ratio on the vertical axis, the better the reconstruction performance; the lower the bit rate bpp on the horizontal axis, the better the compression efficiency. It can be seen that the method proposed by the present invention can achieve the best reconstruction performance with the minimum bit rate.

[0063] The above embodiments are the preferred embodiments of the present application. Those of ordinary skill in the art can also make various transformations or improvements on this basis. Without departing from the general concept of the present application, these transformations or improvements should all fall within the scope of protection required by the present application.

Claims

1. A method for compressing and decompressing images of an underwater net-working robot, characterized in that: include: Step 1, constructing an underwater image compression model based on frequency domain cross-scale fusion; The underwater image compression-decompression model adopts a scale hyper-prior architecture, and includes a first encoder, a first quantization module, a first arithmetic encoder, a first arithmetic decoder and a first decoder from input to output; wherein the first encoder includes a frequency domain cross-scale fusion module, and the first encoder and the first decoder both include a de-redundancy module; At the output end of the first encoder, a second encoder, a second quantization module, a second arithmetic encoder, a second arithmetic decoder and a second decoder are sequentially arranged; The first encoder, the first quantization module, the first arithmetic encoder, the second encoder, the second quantization module, the second arithmetic encoder, and the second decoder together constitute a compression module of the underwater image compression-decompression model; The first arithmetic decoder, the first decoder, the second arithmetic decoder, and the second decoder together constitute a decompression module of the underwater image compression-decompression model; Step 2, using the underwater net image training set to train the constructed underwater image compression-decompression model; Step 3, using the trained compression module, compressing the image collected by the underwater net operation robot to obtain corresponding bit stream data; The trained decompression module is used to reconstruct the underwater image from the compressed bit stream data of the acquired underwater image.

2. The method for compressing and decompressing images of an underwater net-working robot according to claim 1, characterized in that: In the compression module of the underwater image compression-decompression model: The first encoder is used to map the underwater image x collected by the underwater net operation robot to a low-dimensional latent space for feature extraction and output the latent representation y; The second encoder is used to capture the spatial dependencies between different elements in the latent representation y and output a low-dimensional latent variable ; The second quantization module is used to convert the continuous latent variables Perform quantization operations to obtain quantized variables ; The second arithmetic encoder is used to quantize the variables The probability distribution of the bit stream is encoded to obtain the second bit stream ; The first quantization module is used to quantize the continuous potential representation y to obtain a quantized representation ; The second decoder is used to quantize the variables To dynamically estimate the quantized representation Standard Deviation , construct a mean of 0 and a standard deviation of Gaussian probability model of; The first arithmetic encoder is used to quantize the representation according to a Gaussian probability model. Encode to obtain the first bit stream ; Among them, the first bit stream With the second bitstream Together they constitute compressed bit stream data of underwater images; and / or, In the decompression module of the underwater image compression-decompression model: The second arithmetic decoder is used to obtain the second bit stream according to Decoding quantized variables ; The second decoder is based on the quantized variable Dynamic Estimation Quantization Representation Standard Deviation , construct a mean of 0 and a standard deviation of Gaussian probability model of; The first arithmetic decoder converts the first bit stream obtained according to the Gaussian probability model Decoding to quantized representation ; The first decoder then converts the quantized representation Reconstruct the underwater image.

3. The method for compressing and decompressing images of an underwater net-working robot according to claim 1, characterized in that: The underwater net image training set is derived from real deep-sea cage aquaculture scenes, and underwater net images at different depths, environments, and resolutions are selected.

4. The method for compressing and decompressing images of an underwater net-working robot according to claim 1, characterized in that: The frequency domain cross-scale fusion module is implemented by setting a spatial domain path and a frequency domain path and complementing the information between the two paths: First, the input feature map F is proportional to the feature channel dimension. Local feature map and global feature map ; Then in the spatial domain path, the local feature maps are and global feature map After convolution processing, the local output feature map is obtained by fusion ; In addition, in the frequency domain path, on the one hand, the local feature map Convolution processing, on the other hand, for the global feature map Downsampling, convolution and Fourier transform are performed to the frequency domain for processing, and then the feature maps obtained from the two aspects are fused to obtain the global output feature map ; Finally, the local output feature map And the global output feature map By splicing along the feature channel dimension, we can obtain the output feature map M of the frequency domain cross-scale fusion module.

5. The method for compressing and decompressing images of an underwater net-working robot according to claim 4, characterized in that: The steps of the spatial domain path are as follows: transform the local feature map Through the input channel , the output channel is The feature map is obtained by ordinary convolution , the global feature map Through the input channel , the output channel is The feature map is obtained by ordinary convolution , we can get the local output feature map : ; in, Represents the number of channels of the feature map F before the input frequency domain cross-scale fusion module, represents the number of channels of the feature map M output from the frequency domain cross-scale fusion module, and are all hyperparameters, is the ratio of local features in the input channel, is the ratio of local output in the output channel.

6. The method for compressing and decompressing images of an underwater net-working robot according to claim 4, characterized in that: The steps of the frequency domain path are as follows: The local feature map Through the input channel , the output channel is The feature map is obtained by ordinary convolution ;in, Represents the number of channels of the feature map F before the input frequency domain cross-scale fusion module, represents the number of channels of the feature map M output from the frequency domain cross-scale fusion module, and are all hyperparameters, is the ratio of local features in the input channel, is the ratio of local output in the output channel; The global feature map First, down-sample through an average pooling layer, reduce the height and width by half, and then pass through a convolution block including a convolution layer, a BN layer and a ReLU activation function, and then pass through three branches respectively: the first branch first passes the feature map through a fast Fourier transform, extracts the real and imaginary parts of the complex result and splices them, then passes through a convolution layer with a convolution kernel of 1×1, a BN layer and a ReLU activation function, and finally separates the real and imaginary parts again and performs an inverse fast Fourier transform, retaining the real part as the output; the second branch outputs directly without any processing; the third branch first divides the feature map into 4 parts along the height and width respectively, and re-splices it into a feature map along the channel dimension, processes it through the Fourier unit module, and then copies four copies for stacking, and finally outputs a feature map with the same height and width as the input feature map of the branch; then the outputs of the three branches are added and passed through a convolution layer with a convolution kernel of 1×1 to obtain the feature map ; Wherein, the processing method of the first branch is recorded as the Fourier unit module processing; Finally, the feature map obtained by the frequency domain path and Add together to get the global output feature map : 。 7. The method for compressing and decompressing images of an underwater net-working robot according to claim 1, characterized in that: The de-redundancy module is applied after the last convolution layer of the first encoder and before the first convolution layer of the first decoder; the de-redundancy module first copies the input feature map, then initializes a 3×3 convolution layer, applies the convolution layer to some channels of the input feature map, and keeps the other channels unchanged, and performs forward propagation in a slicing manner.

8. A device for compressing images of an underwater net-working robot, characterized in that: It is composed of a compression module trained in the compression and decompression method according to any one of claims 1 to 7; The first encoder is used to map the underwater image x collected by the underwater net operation robot to a low-dimensional latent space for feature extraction and output the latent representation y; The second encoder is used to capture the spatial dependencies between different elements in the latent representation y and output a low-dimensional latent variable z; The second quantization module is used to quantize the continuous latent variable z to obtain the quantized variable ; The second arithmetic encoder is used to quantize the variables The probability distribution of the bit stream is encoded to obtain the second bit stream ; The first quantization module is used to quantize the continuous potential representation y to obtain a quantized representation ; The second decoder is used to quantize the variables To dynamically estimate the quantized representation Standard Deviation , construct a mean of 0 and a standard deviation of Gaussian probability model of; The first arithmetic encoder is used to quantize the representation according to a Gaussian probability model. Encode to obtain the first bit stream ; Among them, the first bit stream With the second bitstream Together they constitute the compressed bit stream data of the underwater image.

9. A decompression device for underwater net-working robot images, characterized in that: It is composed of a decompression module trained in the compression and decompression method according to any one of claims 1 to 7; The second arithmetic decoder is used to obtain the second bit stream according to Decoding quantized variables ; The second decoder is based on the quantized variable Dynamic Estimation Quantization Representation Standard Deviation , construct a mean of 0 and a standard deviation of Gaussian probability model of; The first arithmetic decoder converts the acquired first bit stream into Decoding to quantized representation ; The first decoder then converts the quantized representation Reconstruct the underwater image.

Citation Information

Patent Citations

  • Multi-modal medical image fusion method based on multi-scale codec

    CN116757982A

  • Seismic data reconstruction method based on spatial domain and frequency domain fusion architecture

    CN118033732A

  • Residual error enhanced frequency space mutual learning face super-resolution method

    CN118333860A

  • Underwater image enhancement method of Mama hybrid architecture based on space-frequency fusion

    CN118710507A

  • Space and frequency domain combined super-resolution reconstruction method for arbitrary-scale multi-modal remote sensing image

    CN118799183A