Underwater image man-machine co-friendly compression method based on mixed prior embedding
Through the deep learning framework of mixed prior embeddings and combining internal and external prior information, the problem of low reconstruction quality of human-computer vision tasks in the prior art is solved, and efficient underwater image compression and high-quality reconstruction are achieved.
Patent Information
- Application Number
- CN202510734944.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-04
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2045-06-04
AI Technical Summary
The existing underwater image compression methods are difficult to meet the needs of human-computer visual tasks at the same time, the reconstruction image quality is not high, and the underwater prior knowledge is not effectively utilized.
Underwater image compression method based on hybrid prior embedding is adopted, an underwater image compression network is built through a deep learning framework, combined with internal physical priors and high-quality external implicit prior codebooks, and feature map processing is used to achieve human-computer friendly image compression.
Improves underwater image compression rate and reconstruction image quality, improves the performance of human-computer vision tasks, especially maintaining image clarity and detail at low bit rates.
Smart Images

Figure CN120263983A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image compression methods. Specifically, it relates to an underwater image human-machine co-friendly compression method based on hybrid prior embedding. Background Art
[0002] Underwater image compression aims to efficiently reduce the storage space and transmission bandwidth requirements of underwater image data, while maximizing the retention of key information in the image to meet the actual application needs of underwater imaging devices under limited storage and transmission capabilities. As a key underwater visual information processing technology, underwater image compression has been widely applied in many fields such as marine scientific research, underwater archaeology, marine resource exploration, underwater robot navigation, and marine ecological protection. In recent years, with the rapid development of deep learning, technologies such as convolutional neural networks and Transformers have achieved good results in image compression tasks. However, when facing the joint degradation of underwater imaging and compression, conventional image compression methods still encounter many problems.
[0003] In existing underwater image compression methods, the mainly widely used ones are wavelet-based algorithms, principal component-based algorithms, and CNN-based algorithms. Wavelet-based algorithms utilize the human visual system (HVS) to eliminate visual redundancy. Principal component-based algorithms aim to extract and compress compact principal components. Since underwater images have more low-frequency information than ground images, these methods have improved the compression ratio to a certain extent, but still exhibit blurring at low bitrates. CNN-based algorithms combine CNN and traditional codecs to obtain better reconstruction quality in an end-to-end manner, but their results are only suitable for human viewing, and they do not take into account the importance of machine vision in underwater applications. Currently, there have also been some works that utilize underwater prior knowledge to obtain machine-friendly features to improve analysis performance. However, in the goals of machine vision and human vision, there is still little research on how to correctly introduce underwater priors in the compression framework and how to obtain more compact feature representations, and the compressed images obtained by the above methods cannot guarantee high quality. Therefore, there is an urgent need for an underwater image compression method that can transcend the limitations of existing technologies, achieve simultaneous orientation towards human and machine vision, and improve the quality of reconstructed images. Summary of the Invention
[0004] The technical problem to be solved by the present invention is to efficiently achieve underwater image compression for both human and machine vision tasks and improve the quality of reconstructed images. To overcome the defects of the above prior art (or related art), the present invention provides an underwater image human-machine co-friendly compression method based on hybrid prior embedding.
[0005] The present invention provides an underwater image human-machine co-friendly compression method based on hybrid prior embedding, including the following steps: Step S1: Obtain multiple first original underwater images and preprocess them to divide them into a training set and a test set; Step S2: Obtain multiple second original underwater images, perform human visual preliminary screening and machine vision final screening in sequence, and use the VQ-GAN algorithm for discrete representation learning and high-quality reconstruction to obtain a high-quality external implicit prior codebook; Step S3: Build an underwater image compression network using a deep learning framework. Extract a transmission map as an internal physical prior from the preprocessed first original underwater image through the underwater image compression network. Perform convolutional downsampling on the transmission map and the preprocessed first original underwater image respectively to generate feature maps and feature maps , for the feature maps and feature maps perform feature addition to generate a feature map , for the feature map perform hyperprior analysis to obtain a feature map , perform nearest neighbor search in the high-quality external implicit prior codebook according to the feature map to obtain a feature map , perform convolutional upsampling on the feature map and the feature map based on a bidirectional cross-attention mechanism to obtain a reconstructed image; Step S4: Train the underwater image compression network using the training set according to Step S3 to obtain an underwater image compression network model; Step S5: Test the test set through the underwater image compression network model and output each reconstructed image as the underwater image compression result.
[0006] In the present invention, hybrid prior information of the original underwater image is introduced from two perspectives of internal prior and external prior to form a training set, a test set, and a high-quality external implicit prior codebook. When the underwater image compression network is processing, the internal prior of the training set is used to assist in promoting the features in the first original underwater image to become compact, and the external prior of the high-quality external implicit prior codebook is used to assist in improving the quality of the reconstructed image, thereby improving the image compression rate to enhance the performance of the human-machine vision task. Moreover, in the present invention, a parallel underwater image compression network integrating a bidirectional cross-attention mechanism is designed, which effectively integrates the high-quality codebook characteristics brought by the high-quality external implicit prior codebook and further improves the quality of the reconstructed image.
[0007] In a possible implementation manner, in the step S1, the process of preprocessing each of the first original underwater images includes: Scale the size of each of the first original underwater images to , then perform normalization processing on each of the scaled first original underwater images, normalize the pixel values of all pixel points in the R channel of each of the scaled first original underwater images to a mean of 0.485 and a variance of 0.229, normalize the pixel values of all pixel points in the G channel to a mean of 0.456 and a variance of 0.224, and normalize the pixel values of all pixel points in the B channel to a mean of 0.406 and a variance of 0.225.
[0008] In a possible implementation manner, in step S2, before performing discrete representation learning and high-quality reconstruction on each of the second original underwater images, there is also a preprocessing process, and the preprocessing process includes: Perform initial screening for human vision and final screening for machine vision on each of the second original underwater images in sequence to obtain screened images, and scale the size of each of the screened images to , then perform normalization processing on each of the scaled screened images, normalize the pixel values of all pixel points in the R channel of each of the scaled screened images to a mean of 0.485 and a variance of 0.229, normalize the pixel values of all pixel points in the G channel to a mean of 0.456 and a variance of 0.224, and normalize the pixel values of all pixel points in the B channel to a mean of 0.406 and a variance of 0.225.
[0009] In a possible implementation manner, in step S3, use a deep learning framework to build the underwater image compression network. The underwater image compression network includes a transmission map extraction module, two feature extraction modules, a feature conversion module, a compressor module, a codebook retrieval module, and an output processing module. Receive the preprocessed first original underwater image through the transmission map extraction module and perform feature extraction to obtain the transmission map. Receive the preprocessed first original underwater image through one of the feature extraction modules and perform convolutional downsampling to generate the feature map , receive the transmission map through the other feature extraction module and perform convolutional downsampling to generate the feature map , through the feature conversion module, perform feature addition conversion on the feature map and the feature map to generate the feature map , through the compressor module, perform hyperprior analysis on the feature map to obtain the feature map , through the codebook retrieval module, perform nearest neighbor search in the high-quality external implicit prior codebook according to the feature map to obtain the feature map , through the output processing module, based on the bidirectional cross-attention mechanism, process the feature map and the feature map Perform convolutional upsampling on it to obtain the reconstructed image.
[0010] In a possible implementation, in step S3, the transmission map extraction module includes a depth calculation layer, a background light calculation layer, and a transmission map calculation layer connected in sequence. Through the depth calculation layer, the preprocessed first original underwater image is converted from the RGB space to the YCbCr space, and the horizontal gradient and vertical gradient of the Y channel are calculated. The horizontal gradient and the vertical gradient are combined into a scalar gradient magnitude map. Subsequently, dilation operation, small hole filling operation, and normalization operation are sequentially performed on the scalar gradient magnitude map to obtain a processed gradient magnitude map. According to the processed gradient magnitude map, a preliminary depth map is calculated, and normalization operation and linear regression are performed in combination with the preprocessed first original underwater image to obtain the weights of each color channel, and minimum filtering is calculated for each color channel according to the weights. The minimum values at each pixel position are merged after minimum filtering for each to generate the feature map D; through the background light calculation layer, vectorization operation is performed on the feature map D, and the pixel values in the feature map D are sorted. The pixel value with the largest brightness is selected as the pixel index. After extracting the pixel values at each pixel position based on the pixel index, the average value is taken to generate the feature map A; through the transmission map calculation layer, the color difference value between the preprocessed first original underwater image and the feature map A is calculated and normalized. The maximum value of the normalized color difference value is taken at each pixel position, and median filtering is performed on the maximum value and then normalized to generate the transmission map.
[0011] In a possible implementation, in step S3, the two feature extraction modules are the first feature extraction module and the second feature extraction module. Both the first feature extraction module and the second feature extraction module include a first convolutional downsampling block, a second convolutional downsampling block, a third convolutional downsampling block, and a fourth convolutional downsampling block connected in sequence. The first convolutional downsampling block of the first feature extraction module receives the preprocessed first original underwater image to generate a feature map output, and the second convolutional downsampling block receives the feature map to generate a feature map output, the third convolutional downsampling block receives the feature map to generate a feature map output, and the fourth convolutional downsampling block receives the feature map to generate the feature map output; the first convolutional downsampling block of the second feature extraction module receives the transmission map to generate a feature map output, and the second convolutional downsampling block receives the feature map to generate a feature map Output, the third convolutional downsampling block receives the feature map Generate a feature map Output, the fourth convolutional downsampling block receives the feature map Generate the feature map Output; the feature map has a size of The feature map has a size of The feature map has a size of The feature map has a size of The feature map has a size of The feature map has a size of The feature map has a size of The feature map has a size of .
[0012] In a possible implementation manner, in the step S3, the feature conversion module includes a first-layer feature conversion block and a second-layer feature conversion block that are the same in structure and are connected in sequence. The first-layer feature conversion block receives the feature map and the feature map to perform feature addition conversion to generate a feature map Output, the second-layer feature conversion block receives the feature map to generate a feature map Output, the feature map has a size of The feature map has a size of .
[0013] In a possible implementation manner, in the step S3, the compressor module receives the feature map to perform entropy bottleneck processing and quantization processing, and then uses a checkerboard pattern to perform block processing on the feature map to obtain a plurality of feature blocks, perform conditional entropy coding on each feature block and use noise quantization to process each feature map, and merge the processed feature blocks to obtain an encoded feature map as the feature map .
[0014] In a possible implementation manner, in the step S3, the codebook retrieval module receives the feature map and uses the Euclidean distance as a metric for nearest neighbor search, and the feature map Map to find the closest codeword in the high-quality external implicit prior codebook And transform to obtain the feature map .
[0015] In a possible implementation, in step S3, the output processing module includes a first branch and a second branch. The first branch includes a first-layer convolutional upsampling block, a second-layer convolutional upsampling block, a third-layer convolutional upsampling block, and a fourth-layer convolutional upsampling block connected in sequence. The second branch includes a first-layer bidirectional cross-attention module, a first-layer convolutional upsampling block, a second-layer bidirectional cross-attention module, a second-layer convolutional upsampling block, a third-layer bidirectional cross-attention module, a third-layer convolutional upsampling block, a fourth-layer bidirectional cross-attention module, and a fourth-layer convolutional upsampling block connected in sequence. The first-layer convolutional upsampling block in the first branch receives the feature map And generates a feature map Output. The second-layer convolutional upsampling block receives the feature map And generates a feature map Output. The third-layer convolutional upsampling block receives the feature map And generates a feature map Output. The first-layer bidirectional cross-attention module in the second branch receives the feature map And the feature map And generates a feature map through the first-layer convolutional upsampling block . The second-layer bidirectional cross-attention module receives the feature map And the feature map And generates a feature map through the second-layer convolutional upsampling block . The third-layer bidirectional cross-attention module receives the feature map And the feature map And generates a feature map through the third-layer convolutional upsampling block . The fourth-layer bidirectional cross-attention module receives the feature map And the feature map And generates the reconstructed image through the fourth-layer convolutional upsampling block. The size of the feature map Is . The size of the feature map Is . The size of the feature map Is . The size of the feature map Is ; The size of the feature map Is . The size of the feature map Is . The size of the feature map The size of is The size of . Brief Description of the Drawings
[0016] Figure 1 is the flowchart of the steps of the present invention; Figure 2 is the schematic diagram of the framework of the underwater image compression network of the present invention; Figure 3 is the schematic diagram of the framework of the transmission map extraction module of the present invention; Figure 4 is the schematic diagram of the framework of the feature conversion module of the present invention; Figure 5 is the schematic diagram of the framework of the compressor module of the present invention; Figure 6 is the schematic diagram of the framework of the bidirectional cross-attention module of the present invention; Figure 7 is the schematic diagram of the visual effect and machine task performance of some data of the method of the present invention and the existing image compression method on the SUIM, RUOD, and USOD test sets. Detailed Embodiments
[0017] First of all, those skilled in the art should understand that these embodiments are only used to explain the technical principles of the embodiments of the present invention and are not intended to limit the protection scope of the embodiments of the present invention. Those skilled in the art can adjust them according to needs to adapt to specific application scenarios.
[0018] The present invention will be further described in detail below with reference to the drawings and specific embodiments.
[0019] Referring to Figure 1 , the embodiments of the present invention disclose an underwater image human-machine co-friendly compression method based on hybrid prior embedding, including the following steps: Step S1, obtaining multiple first original underwater images and preprocessing them into a training set and a test set; Step S2, obtaining multiple second original underwater images, successively performing human visual preliminary screening and machine visual final screening, and using the VQ-GAN algorithm for discrete representation learning and high-quality reconstruction to obtain a high-quality external implicit prior codebook; Step S3, using a deep learning framework to build an underwater image compression network, extracting a transmission map as an internal physical prior from the preprocessed first original underwater image through the underwater image compression network, and respectively performing convolutional downsampling on the transmission map and the preprocessed first original underwater image to generate a feature map and a feature map , for the feature map and the feature map Generate a feature map by adding features , for the feature map Perform a hyperprior analysis to obtain a feature map , according to the feature map Perform a nearest neighbor search in a high-quality external implicit prior codebook to obtain a feature map , based on a bidirectional cross-attention mechanism for the feature map and the feature map Perform convolutional upsampling to obtain a reconstructed image; Step S4, train an underwater image compression network using the training set according to Step S3 to obtain an underwater image compression network model; Step S5, test the test set through the underwater image compression network model and output each reconstructed image as the underwater image compression result.
[0020] In the embodiment of the present invention, in Step S1, a dataset is selected or constructed, and each first original underwater image has an underwater object with a relatively complex morphological structure; then, each first original underwater image in the dataset is preprocessed including a cropping operation and a normalization operation, so that the size of the preprocessed first original underwater image is ; then all the preprocessed first original underwater images are divided into a training set and a test set, and the image distribution in the training set is the same as that in the test set. The proportion of images with various categories such as divers, fish, corals, etc. in the training set is uniform, and it is also this proportion in the test set; among them, in this embodiment , .
[0021] In the embodiment of the present invention, specifically in operation, the process of preprocessing a first original underwater image in Step S1 is as follows: First, use the existing technology to crop the size of the first original underwater image to ; secondly, perform a normalization process on the scaled first original underwater image, normalize the pixel values of all pixel points in the R channel of the scaled first original underwater image to a mean of 0.485 and a variance of 0.229, normalize the pixel values of all pixel points in the G channel to a mean of 0.456 and a variance of 0.224, and normalize the pixel values of all pixel points in the B channel to a mean of 0.406 and a variance of 0.225; since the input of a deep neural network usually has a fixed form, it is necessary to preprocess the first original underwater image.
[0022] In the embodiment of the present invention, in step S2, a dataset friendly to human vision and machine vision tasks is constructed. The dataset includes various underwater scenes, and each second original underwater image has high resolution and excellent performance in various machine task algorithms. Then, each second original underwater image in the dataset is preprocessed including cropping operation and normalization operation, so that the size of the preprocessed second original underwater image is , the pixel values of all pixel points in the R channel of the scaled second original underwater image are normalized to a mean of 0.485 and a variance of 0.229, the pixel values of all pixel points in the G channel are normalized to a mean of 0.456 and a variance of 0.224, and the pixel values of all pixel points in the B channel are normalized to a mean of 0.406 and a variance of 0.225; a well-trained VQ codebook is constructed. The VQ codebook is a high-quality external implicit prior codebook generated by the prior art VQ-GAN algorithm and is trained on a large corpus of underwater images friendly to human-machine vision.
[0023] In the embodiment of the present invention, in specific operation, the process of introducing a high-quality external implicit prior codebook as a credible and reliable external underwater feature library in step S2 is as follows: First, find underwater datasets SUIM, RUOD, and USOD that contain multiple underwater scenes and objects; Second, perform a preliminary screening for human vision on the second original underwater images of the three datasets. Those with a BRISQUE index higher than 50 are excluded; Third, perform a final screening for machine vision on the images of the three datasets. Set the top fifty percent according to the machine task performance corresponding to each dataset. The intersection over union (IOU) is used as the evaluation index for object detection, the Sα is used as the evaluation index for semantic segmentation, and the average precision (AP) is used as the evaluation index for salient object detection; Among them, the IOU is a key index for measuring the positioning accuracy in object detection; the Sα is an index for evaluating the segmentation accuracy of salient objects in semantic segmentation; the AP is an index for comprehensively evaluating the precision and recall rate of the model in salient object detection; Subsequently, the size of the second original underwater image in the constructed new dataset is cropped to ; Second, perform normalization processing on the scaled second original underwater image. Finally, transfer the data to the VQGAN network. By combining the encoder, decoder, and discriminator, the discrete representation learning and high-quality reconstruction of the second original underwater image are realized. During the training process, vector quantization ensures the discretization of features, while adversarial training improves the perceptual quality of the generated images. By optimizing the combination of reconstruction loss, vector quantization loss, and adversarial loss, the VQGAN network can generate high-quality images and learn an efficient discrete codebook, that is, a credible and reliable external underwater feature library.
[0024] In an embodiment of the present invention, in step S3, a deep neural network is built using a deep learning framework as an underwater image compression network based on hybrid prior assistance, as Figure 2 shown. The underwater image compression network includes a transmission map extraction module, a feature extraction module, a feature conversion module, a compressor module, a codebook retrieval module, and an output processing module.
[0025] In an embodiment of the present invention, as Figure 3 shown, the transmission map extraction module mainly includes three layers: a depth calculation layer, a background light calculation layer, and a transmission map calculation layer. The input end of the first layer, the depth calculation layer, of the transmission map extraction module serves as the input end of the transmission map extraction module to receive an RGB image of size , that is, the first original underwater image. The rgb_to_ycbcr function is used to convert the first original underwater image from the RGB space to the YCbCr space. The cv2.Sobel function is used to calculate the horizontal gradient and vertical gradient of the Y channel, and the horizontal gradient and vertical gradient are combined into a scalar gradient magnitude map. Subsequently, the cv2.dilate function is used to perform a dilation operation on the scalar gradient magnitude map. The processed gradient map is obtained by filling the small holes in the scalar gradient magnitude map and normalizing the scalar gradient magnitude map. Then, a preliminary depth map is calculated. Both the preliminary depth map and the scalar gradient magnitude map are vectorized and linearly regressed to obtain the weights of each color channel, and the local minimum filtering is calculated for each color channel respectively. After merging, the minimum value at each pixel position is taken, and the feature map output from the output end of the first layer depth calculation layer is denoted as D. The input end of the second layer, the background light calculation layer, of the transmission map extraction module receives the feature map D output from the first layer depth calculation layer. The feature map D is vectorized, the pixel values of the feature map D are sorted, and the pixel index with the highest brightness is selected. After extracting the corresponding pixel value from the flattened image and taking the average, the atmospheric light value is obtained, and the feature map output from the output end of the second layer background light calculation layer is denoted as A. The input end of the third layer, the transmission map calculation layer, of the transmission map extraction module receives the feature map A output from the output end of the second layer background light calculation layer; the color difference between the first original underwater image and the feature map A is calculated and normalized, and the maximum value of the normalized difference map is taken at each pixel position. After median filtering the maximum value and normalizing it, the feature map output from the output end of the third layer transmission map calculation layer is denoted as T-map. The output end of the third layer of the transmission map extraction module serves as the output end of the transmission map extraction module.
[0026] In an embodiment of the present invention, the feature extraction module is a ResNet50 backbone network with four layers and four convolutional downsampling blocks. The four convolutional downsampling blocks are sequentially connected and inserted at intervals in the four-layer ResNet50 backbone network. The input end of the first layer ResNet50 backbone network serves as the input end of the feature extraction module to receive an RGB image of size The RGB image, i.e., the first original underwater image or the transmission map extracted after the transmission map extraction module, has a size of The output end of the fourth-layer ResNet50 backbone network serves as the output end of the feature extraction module after passing through the convolutional downsampling block.
[0027] In the embodiment of the present invention, for the RGB image, the feature map output by the output end of the first convolutional downsampling block is denoted as The second convolutional downsampling block receives the feature map output by the output end of the first convolutional downsampling block The feature map output by the output end of the second convolutional downsampling block is denoted as The third convolutional downsampling block receives the feature map output by the output end of the second convolutional downsampling block The feature map output by the output end of the third convolutional downsampling block is denoted as The fourth convolutional downsampling block receives the feature map output by the output end of the third convolutional downsampling block The feature map output by the output end of the fourth convolutional downsampling block is denoted as ; among them, the size of the feature map is The size of the feature map is The size of the feature map is The size of the feature map is .
[0028] In the embodiment of the present invention, for the transmission map, the feature map output by the output end of the first convolutional downsampling block is denoted as The second convolutional downsampling block receives the feature map output by the output end of the first convolutional downsampling block The feature map output by the output end of the second convolutional downsampling block is denoted as The third convolutional downsampling block receives the feature map output by the output end of the second convolutional downsampling block The feature map output by the output end of the third convolutional downsampling block is denoted as The fourth convolutional downsampling block receives the feature map output by the output end of the third convolutional downsampling block The feature map output by the output end of the fourth convolutional downsampling block is denoted as ; among them, the size of the feature map is The size of the feature map is The size of the feature map is The size of the feature map is .
[0029] In the embodiments of the present invention, each convolutional downsampling module includes 3 residual blocks and 1 convolutional block of one layer of the ResNet50 backbone network. One convolutional block is followed by 3 residual blocks. Each residual block has the same structure and includes a convolutional layer, a ReLU activation function layer, a convolutional layer, a ReLU activation function layer, and a convolutional layer connected in sequence. The convolutional kernel size of the first convolutional layer is 1, the stride size is 1, the padding is 0, the number of input channels is 192, and the number of output channels is 96. The convolutional kernel size of the second convolutional layer is 3, the stride size is 1, the padding is 1, the number of input channels is 96, and the number of output channels is 96. The convolutional kernel size of the third convolutional layer is 1, the padding is 1, the number of input channels is 96, and the number of output channels is 192. Each convolutional downsampling module contains 4 convolutional blocks. The convolutional kernel size of the first convolutional block is 5, the stride size is 2, the padding is 2, the number of input channels is 3, and the number of output channels is 192. The convolutional kernel size of the second convolutional block is 5, the stride size is 2, the padding is 2, the number of input channels is 192, and the number of output channels is 192. The convolutional kernel size of the third convolutional block is 5, the stride size is 2, the padding is 2, the number of input channels is 192, and the number of output channels is 192. The convolutional kernel size of the fourth convolutional block is 5, the stride size is 2, the padding is 2, the number of input channels is 192, and the number of output channels is 320.
[0030] In the embodiments of the present invention, the feature conversion module is mainly composed of two feature conversion blocks with the same structure. The two layers in the feature conversion module are connected in sequence, as Figure 4 shown. The input end of the first-layer feature conversion block receives the added feature of the feature map output by the output end of the feature extraction module of the original image branch and the feature map output by the output end of the feature extraction module of the transmission image branch . The feature map output by the output end of the first-layer feature conversion block is denoted as ; The input end of the second-layer feature conversion block receives the feature map output by the output end of the first-layer feature conversion block , and the feature map output by the output end of the second-layer feature conversion block is denoted as .
[0031] In the embodiments of the present invention, each layer of the feature conversion block is a parallel structure. One branch includes a pooling layer, a convolutional layer, a batch normalization layer, a ReLU activation function layer, a convolutional layer, and a batch normalization layer connected in sequence. The other branch includes a convolutional layer, a batch normalization layer, a ReLU activation function layer, a convolutional layer, and a batch normalization layer connected in sequence. After adding the two branches and performing a Sigmoid activation function operation, they are respectively combined with the feature map output by the output end of the fourth layer of the feature extraction module of the original image branch received by the input end of the first-layer feature conversion block and the feature map output from the output end of the fourth layer of the feature extraction module of the transmission map branch After multiplying and then adding them together, the feature map output from the output end of the i-th layer feature conversion block is denoted as , where the convolution kernel size of the first convolutional layer is 1, the stride size is 1, the padding is 0, the number of input channels is 320, and the number of output channels is 80; the convolution kernel size of the second convolutional layer is 1, the stride size is 1, the padding is 0, the number of input channels is 80, and the number of output channels is 320.
[0032] In the embodiment of the present invention, as Figure 5 shown, the compressor module mainly consists of a spatio-channel context adaptive module that first applies a spatial context model to eliminate spatial redundancy and then eliminates channel redundancy through a channel conditional model, an arithmetic encoder module that maps the quantized encoded symbols to a real number interval and converts them into a binary bitstream, an arithmetic decoder module that recovers the symbol sequence from the bitstream and decodes it using the same probability model, a hyperprior analyzer module that extracts the hyperprior information of the latent space and optimizes the entropy coding process, and a hyperprior synthesizer module that decodes the hyperprior information and helps the entropy decoder recover the latent space symbols. The input end of the compressor module receives the feature map output from the output end of the feature conversion module , performs entropy bottleneck processing and quantization on the input feature map , uses a checkerboard pattern to perform block processing on the feature map , performs conditional entropy coding on each feature block, calculates its probability, uses noise quantization to process the feature block, and combines the processed feature blocks to obtain the final encoded feature map, that is, the feature map output from the output end of the compressor module Among them, the probability of entropy coding reflects the change in the size of bpp and is used for the construction of the loss function. During testing, since differentiability is not considered, the bpp value can be directly calculated.
[0033] In the embodiment of the present invention, the hyperprior analyzer module includes a convolutional layer, a ReLU activation function layer, a convolutional layer, a ReLU activation function layer, and a convolutional layer connected in sequence. The convolutional kernel size of the first convolutional layer is 3, the stride size is 1, the padding is 1, the number of input channels is 320, and the number of output channels is 192. The convolutional kernel size of the second convolutional layer is 1, the stride size is 1, the padding is 0, the number of input channels is 192, and the number of output channels is 192. The hyperprior synthesizer module includes a transposed convolutional layer, a ReLU activation function layer, a transposed convolutional layer, a ReLU activation function layer, and a transposed convolutional layer connected in sequence. The convolutional kernel size in the first transposed convolutional layer is 1, the stride size is 1, the padding is 0, the number of input channels is 192, and the number of output channels is 192. The convolutional kernel size in the second transposed convolutional layer is 1, the stride size is 1, the padding is 0, the number of input channels is 192, and the number of output channels is 288. The convolutional kernel size in the third transposed convolutional layer is 3, the stride size is 1, the padding is 1, the number of input channels is 288, and the number of output channels is 640.
[0034] In the embodiment of the present invention, the codebook retrieval module is mainly composed of the vector quantization (Vector Quantization) technology of VQGAN, and the feature map output from the output end of the compressor module is used as the input of the codebook retrieval module and uses the nearest neighbor search to map to the closest codeword in the codebook The obtained feature map output from the output end is denoted as .
[0035] In the embodiment of the present application, as Figure 2 shown, the output processing module is mainly composed of a bidirectional cross-attention module and a convolutional upsampling module, and is divided into two branches. The first input end of the output processing module receives the feature map output from the output end of the compressor module , and the second input end receives the feature map output from the output end of the codebook retrieval module . In the first branch, the first convolutional upsampling block receives the feature map output from the output end of the codebook retrieval module , and the feature map output from the output end of the first convolutional upsampling block is denoted as . The second convolutional upsampling block receives the feature map output from the output end of the codebook retrieval module , and the feature map output from the output end of the second convolutional upsampling block is denoted as . The third convolutional upsampling block receives the feature map output from the output end of the codebook retrieval module , and the feature map output from the output end of the third convolutional upsampling block is denoted as ; In the second branch, the first bidirectional cross-attention module receives the feature map output from the output end of the compressor module The feature map output from the codebook retrieval module After passing through the first convolutional upsampling block, the obtained feature map is denoted as The second bidirectional cross-attention module receives the feature map output from the output end of the first convolutional upsampling block and the feature map After passing through the second convolutional upsampling block, the obtained feature map is denoted as The third bidirectional cross-attention module receives the feature map output from the output end of the second convolutional upsampling block and the feature map After passing through the third convolutional upsampling block, the obtained feature map is denoted as The fourth bidirectional cross-attention module receives the feature map output from the output end of the third convolutional upsampling block and the feature map After passing through the fourth convolutional upsampling block, the obtained image is the reconstructed image. Among them, the size of the feature map is The size of the feature map is The size of the feature map is The size of the feature map is ; The size of the feature map is The size of the feature map is The size of the feature map is The size of the feature map is .
[0036] In the embodiment of the present invention, for the first branch, the convolutional kernel size of the convolutional layer in the first convolutional block is 3, the stride size is 1, the padding is 1, the number of input channels is 320, and the number of output channels is 256; the convolutional kernel size of the convolutional layer in the second convolutional block is 1, the stride size is 1, the padding is 0, the number of input channels is 256, and the number of output channels is 128; the convolutional kernel size of the convolutional layer in the third convolutional block is 1, the stride size is 1, the padding is 0, the number of input channels is 128, and the number of output channels is 64; the convolutional kernel size of the convolutional layer in the fourth convolutional block is 1, the stride size is 1, the padding is 0, the number of input channels is 64, and the number of output channels is 3. For the second branch, the convolutional kernel size of the convolutional layer in the first convolutional block is 3, the stride size is 1, the padding is 1, the number of input channels is 320, and the number of output channels is 256; the convolutional kernel size of the convolutional layer in the second convolutional block is 1, the stride size is 1, the padding is 0, the number of input channels is 256, and the number of output channels is 128; the convolutional kernel size of the convolutional layer in the third convolutional block is 1, the stride size is 1, the padding is 0, the number of input channels is 128, and the number of output channels is 64.
[0037] In the embodiment of the present invention, as Figure 6 shown, the bidirectional cross-attention module in the output processing module is mainly composed of a convolutional layer, a normalization layer, a group normalization layer, an activation function, and a deformable convolutional block. The input end of the bidirectional cross-attention module is connected to the feature map and , , and enters four branches. They enter two branches respectively. The first branch has only one convolutional layer, and the second branch has a convolutional layer and a normalization layer in sequence. Among them, the convolutional kernel size of the first convolutional layer is 3, the stride size is 1, the padding is 1, the number of input channels is 256, and the number of output channels is 256; the convolutional kernel size of the second convolutional layer is 1, the stride size is 1, the padding is 0, the number of input channels is 256, and the number of output channels is 256. They enter two branches respectively. The first branch has only one convolutional layer, and the second branch has a convolutional layer and a normalization layer in sequence. Among them, the convolutional kernel size of the first convolutional layer is 3, the stride size is 1, the padding is 1, the number of input channels is 256, and the number of output channels is 256; the convolutional kernel size of the second convolutional layer is 1, the stride size is 1, the padding is 0, the number of input channels is 256, and the number of output channels is 256. and After the branches passing through the normalization layer are multiplied with each other and then multiplied with their respective other branches, and after performing cross-adding operations, channel concatenation is carried out, and then it enters a convolutional layer, a group normalization layer, an activation function layer, and a convolutional layer connected in sequence. The convolutional kernel size of the first convolutional layer is 3, the stride size is 1, the padding is 1, the number of input channels is 512, and the number of output channels is 256. The convolutional kernel size of the second convolutional layer is 3, the stride size is 1, the padding is 1, the number of input channels is 256, and the number of output channels is 256. After interpolating with the offset value of the previous layer, channel concatenation is performed, and then it enters a convolutional layer, a group normalization layer, and an activation function layer connected in sequence to obtain the offset value of this layer. The convolutional kernel size of the convolutional layer is 3, the stride size is 1, the padding is 1, the number of input channels is 512, and the number of output channels is 256; Finally, it enters a convolutional layer and a deformable convolutional layer connected in sequence, and is concatenated with After channel concatenation, it enters a residual block of a ResNet50 backbone network to obtain the feature map output at the output end of the bidirectional cross-attention module. Among them, the convolutional kernel size of the convolutional layer is 3, the stride size is 1, the padding is 1, the number of input channels is 512, and the number of output channels is 256. The convolutional kernel size of the deformable convolutional layer is 3, the stride size is 1, the padding is 1, the number of input channels is 256, and the number of output channels is 256; Among them, the ResNet50 backbone network and the VQGAN network are both existing structural frameworks, and their network structures have been made public. For example, the ResNet50 backbone network is recorded in the reference K. He, X. Zhang, S. Ren and J. Sun, "Deep Residual Learning for Image Recognition," 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pp. 770-778, 2016. ("Image Recognition Based on Deep Residual Learning"), and the VQGAN network is recorded in Patrick Esser, Robin Rombach, Bjorn Ommer, et al. "Taming Transformers for High-Resolution Image Synthesis." (CVPR), 2021, pp. 12873-12883 (VQGAN structure). Here, the feature addition operation and the feature multiplication operation are conventional operations in the convolutional neural network.
[0038] In the embodiment of the present invention, in step S4, the underwater image compression network based on the hybrid prior assistance is trained using the training set. After each round of training, the underwater image compression network based on the hybrid prior assistance outputs the reconstructed images corresponding to each first original underwater image in the training set, which are correspondingly denoted as , and then the loss of the underwater image compression network is calculated, denoted as , ; where , represents the bitstream size, , represents the first original underwater image corresponding to each reconstructed image in the training set, represents the mean square error function, represents the L1 loss function of the original degraded features and the high-quality codebook features in the codebook retrieval module, represents the weight coefficient, the initial learning rate of the Adam optimizer is , and the weight decay rate is set to 0.1 / 3800 rounds to update the network parameters, and the batch size is 16; in this embodiment, is taken.
[0039] In the embodiment of the present invention, the underwater image compression network model based on the hybrid prior assistance is trained for a total of 2000 rounds according to the process of step S4.
[0040] In the embodiment of the present invention, in order to further verify the feasibility and effectiveness of the method of the present invention, the following experiments are carried out on the method of the present invention: In the experiment, existing data sets were directly selected for training and testing. The data sets include the SUIM, RUOD, and USOD data sets. The above three data sets are all publicly available data sets and can be directly obtained through the prior art. Training is carried out on the training set of SUIM, and testing is carried out on the test set of DIS5K and the RUOD and USOD data sets, as Figure 7The following figure shows the visual effects and machine task performance diagrams of some data of the method of the present invention and existing image compression methods on the SUIM, RUOD, and USOD test sets. Table 1 below gives the quantitative comparison of the method of the present invention and existing image compression methods (Original, JPEG2000, BPG, VVC, Minnen2018, Cheng2020, Zou2022, ELIC2022) in terms of machine task evaluation metrics (RI, FV, WR, RO, HD, mIoU, mF, RA, mAP, Sα, Fβ) on the datasets of SUIM, RUOD, and USOD. The above existing image compression methods and machine task evaluation metrics are all commonly used methods and evaluation metrics in the field and will not be defined or explained additionally. All the values shown in Table 1 below are the average values of the evaluation metrics obtained after testing on the datasets of SUIM, RUOD, and USOD. Taking 0.05bpp as an example, the machine task evaluation metrics RI, FV, WR, RO, HD, mIoU, mF are obtained by testing at 0.057, 0.053, 0.052, 0.058, 0.056, 0.054bbp. The machine task evaluation metrics RA, mAP are obtained by testing at 0.057, 0.053, 0.051, 0.055, 0.058, 0.059, 0.056bbp. The machine task evaluation metrics Sα, Fβ are obtained by testing at 0.057, 0.053, 0.05, 0.058, 0.055, 0.052bbp. Table 1 is as follows: Table 1 Quantitative comparison table of the method of the present invention and existing image compression methods (Original, JPEG2000, BPG, VVC, Minnen2018, Cheng2020, Zou2022, ELIC2022) in terms of machine task evaluation metrics (RI, FV, WR, RO, HD, mIoU, mF, RA, mAP, Sα, Fβ) on the datasets of SUIM, RUOD, and USOD ; Table 2 below gives the quantitative comparison of the method of the present invention and existing image compression methods (JPEG2000, BPG, VVC, Minnen2018, Cheng2020, Zou2022) in terms of human visual evaluation metrics (PSNR, MS-SSIM, PI, BRISQUE, LPIPS) on the SUIM test set. The above existing image compression methods and human visual evaluation metrics are all commonly used methods and evaluation metrics in the field and will not be defined or explained additionally. Taking 0.05bpp, 0.08bpp, and 0.12bpp as examples: Table 2 Quantitative comparison table of the method of the present invention and existing image compression methods (JPEG2000, BPG, VVC, Minnen2018, Cheng2020, Zou2022) in terms of human visual evaluation metrics (PSNR, MS-SSIM, PI, BRISQUE, LPIPS) on the test set of SUIM ; For machine vision, the method of the present invention realizes semantic segmentation, object detection, and saliency detection to measure the analysis performance of the method of the present invention. For semantic segmentation, the mean intersection-over-union (mIOU) and slice coefficient (F) are used to evaluate the target boundary localization performance. mIOU calculates the average ratio of the intersection area to the union area of the same classes between semantic maps. F quantifies the correctness of the predicted pixel labels compared to the ground truth, and it is calculated using precision (P) and recall (R) as F = 2×P×R / (P+R). For object detection, recognition accuracy (RA) and mean average precision (mAP) are used to measure the detection performance. RA is the proportion of correctly detected samples, and mAP calculates the average detection precision of different objects. For saliency detection, the average F-measure (Fβ) and S-measure (Sα) are used to evaluate the saliency object localization performance. Fβ is calculated from the weighted harmonic mean of precision and recall, and Sα evaluates the region-aware and object-aware structural similarity between saliency maps; For human vision, the method of the present invention uses PI, BRISQUE, and LPIPS to evaluate the perceptual quality of the reconstructed image and its similarity to the original image. Both PI and BRISQUE are no-reference image quality metrics for natural images and are superior to the traditional PSNR and MS-SSIM metrics in evaluating visual perceptual quality. In addition, the full-reference image quality evaluation metric LPIPS is used to measure the similarity between the reconstructed image and the original image. For these metrics, the lower the score, the better the quality. It can be seen that the method of the present invention has better performance and picture compression effect compared to existing image compression methods.
[0041] In the description of the present invention, the descriptions referring to terms such as "one embodiment", "some embodiments", "in this embodiment", "specific examples", or "some examples" etc. mean that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in a suitable manner in any one or more embodiments or examples. In addition, without contradiction, those skilled in the art can combine and combine the different embodiments or examples described in this specification and the features of different embodiments or examples.
[0042] As described above, it is only the specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any changes or substitutions that can be easily thought of by those skilled in the art within the technical scope disclosed by the present invention should be covered within the protection scope of the present invention. Therefore, the protection scope of the present invention shall be subject to the protection scope of the claims.
Claims
1. An underwater image human-machine co-friendly compression method based on hybrid prior embedding, characterized in that It includes the following steps: Step S1: Obtain multiple first original underwater images and preprocess them to divide them into a training set and a test set; Step S2: Obtain multiple second original underwater images, perform human visual preliminary screening and machine vision final screening in sequence, and use the VQ-GAN algorithm to perform discrete representation learning and high-quality reconstruction to obtain a high-quality external implicit prior codebook; Step S3, build an underwater image compression network using a deep learning framework, and extract a transmission map as an internal physical prior from the preprocessed first original underwater image through the underwater image compression network. Perform convolutional downsampling on the transmission map and the preprocessed first original underwater image respectively to generate feature maps and feature maps , for the feature maps and feature maps perform feature addition to generate a feature map , for the feature map perform hyperprior analysis to obtain a feature map , according to the feature map perform a nearest neighbor search in the high-quality external implicit prior codebook to obtain a feature map , based on the bidirectional cross-attention mechanism, perform convolutional upsampling on the feature map and the feature map to obtain a reconstructed image; Step S4: Train an underwater image compression network using the training set according to Step S3 to obtain an underwater image compression network model; Step S5: Test the test set through the underwater image compression network model and output each reconstructed image as the underwater image compression result.
2. The underwater image human-machine co-friendly compression method based on hybrid prior embedding according to claim 1, wherein, In the said Step S1, the process of preprocessing each of the first original underwater images includes: Scale the size of each of the first original underwater images to , and then perform normalization processing on each of the scaled first original underwater images. Normalize the pixel values of all pixel points in the R channel of each of the scaled first original underwater images to a mean of 0.485 and a variance of 0.229, normalize the pixel values of all pixel points in the G channel to a mean of 0.456 and a variance of 0.224, and normalize the pixel values of all pixel points in the B channel to a mean of 0.406 and a variance of 0.
225.
3. The underwater image human-computer co-friendly compression method based on hybrid prior embedding according to claim 1, characterized in that, In the said Step S2, before performing discrete representation learning and high-quality reconstruction on each of the second original underwater images, there is also a preprocessing process, and the preprocessing process includes: Perform initial screening for human vision and final screening for machine vision on each of the second original underwater images in sequence to obtain the screened images, and scale the size of each of the screened images to , and then perform normalization processing on each of the scaled screened images, normalizing the pixel values of all pixel points in the R channel of each of the scaled screened images to a mean of 0.485 and a variance of 0.229, normalizing the pixel values of all pixel points in the G channel to a mean of 0.456 and a variance of 0.224, and normalizing the pixel values of all pixel points in the B channel to a mean of 0.406 and a variance of 0.
225.
4. The underwater image human-machine co-friendly compression method based on hybrid prior embedding according to claim 1, wherein, In step S3, the underwater image compression network is built using a deep learning framework. The underwater image compression network includes a transmission map extraction module, two feature extraction modules, a feature conversion module, a compressor module, a codebook retrieval module, and an output processing module. The preprocessed first original underwater image is received by the transmission map extraction module and feature extraction is performed to obtain the transmission map. The preprocessed first original underwater image is received by one of the feature extraction modules for convolutional downsampling to generate the feature map , and the transmission map is received by the other feature extraction module for convolutional downsampling to generate the feature map . The feature map and the feature map are subjected to feature addition conversion by the feature conversion module to generate the feature map . The feature map is subjected to hyperprior analysis by the compressor module to obtain the feature map . The codebook retrieval module performs a nearest neighbor search in the high-quality external implicit prior codebook based on the feature map to obtain the feature map . The output processing module performs convolutional upsampling on the feature map and the feature map based on a bidirectional cross-attention mechanism to obtain the reconstructed image.
5. The underwater image human-machine co-friendly compression method based on hybrid prior embedding according to claim 4, wherein, In the said Step S3, the transmission map extraction module includes a depth calculation layer, a background light calculation layer, and a transmission map calculation layer connected in sequence. Through the depth calculation layer, the preprocessed first original underwater image is converted from the RGB space to the YCbCr space, and the horizontal gradient and vertical gradient of the Y channel are calculated. The horizontal gradient and the vertical gradient are combined into a scalar gradient magnitude map. Subsequently, dilation operation, small hole filling operation, and normalization operation are performed on the scalar gradient magnitude map in sequence to obtain a processed gradient magnitude map. According to the processed gradient magnitude map, a preliminary depth map is calculated and combined with the preprocessed first original underwater image to perform normalization operation and linear regression to obtain the weights of each color channel, and minimum filtering is calculated for each color channel according to each of the weights. The minimum values at each pixel point position are taken after combining each of the minimum filtering results to generate a feature map D; through the background light calculation layer, vectorization operation is performed on the feature map D and the pixel values in the feature map D are sorted, the pixel value with the largest brightness is selected as the pixel index, and the average value is taken after extracting the pixel values at each pixel point position based on the pixel index to generate a feature map A; Calculate and normalize the color difference value between the preprocessed first original underwater image and the feature map A through the transmission map calculation layer, take the maximum value of the normalized color difference value at each pixel point position, perform median filtering on the maximum value and then perform normalization to generate the transmission map.
6. The underwater image human-machine co-friendly compression method based on hybrid prior embedding according to claim 4, characterized in that, In step S3, the two feature extraction modules are the first feature extraction module and the second feature extraction module. Both the first feature extraction module and the second feature extraction module include a first convolutional downsampling block, a second convolutional downsampling block, a third convolutional downsampling block, and a fourth convolutional downsampling block connected in sequence. The first convolutional downsampling block of the first feature extraction module receives the preprocessed first original underwater image to generate a feature map output, and the second convolutional downsampling block receives the feature map to generate a feature map output. The third convolutional downsampling block receives the feature map to generate a feature map output. The fourth convolutional downsampling block receives the feature map to generate the feature map output; the first convolutional downsampling block of the second feature extraction module receives the transmission map to generate a feature map output, and the second convolutional downsampling block receives the feature map to generate a feature map output. The third convolutional downsampling block receives the feature map to generate a feature map output. The fourth convolutional downsampling block receives the feature map to generate the feature map output; the size of the feature map is , the size of the feature map is , the size of the feature map is , the size of the feature map is , the size of the feature map is , the size of the feature map is , the size of the feature map is , the size of the feature map is .
7. The underwater image human-machine co-friendly compression method based on hybrid prior embedding according to claim 4, wherein In the step S3, the feature conversion module includes a first-layer feature conversion block and a second-layer feature conversion block which have the same structure and are connected in sequence. The first-layer feature conversion block receives the feature map and the feature map to perform feature addition conversion to generate a feature map for output. The second-layer feature conversion block receives the feature map to generate a feature map for output. The size of the feature map is and the size of the feature map is .
8. The underwater image human-machine co-friendly compression method based on hybrid prior embedding according to claim 4, wherein In the step S3, the compressor module receives the feature map to perform entropy bottleneck processing and quantization processing, and then uses a checkerboard pattern for the feature map to perform block processing to obtain a plurality of feature blocks, perform conditional entropy coding on each of the feature blocks, and use noise quantization to process each of the feature maps, and merge the processed feature blocks to obtain an encoded feature map as the feature map .
9. The underwater image human-machine co-friendly compression method based on hybrid prior embedding according to claim 4, characterized in that, In the step S3, the feature map is received by the codebook retrieval module and the Euclidean distance is used as the metric for nearest neighbor search to map the feature map to the high-quality external implicit prior codebook to find the closest codeword and the feature map is obtained through conversion .
10. The underwater image human-machine co-friendly compression method based on hybrid prior embedding according to claim 4, wherein In the step S3, the output processing module includes a first branch and a second branch. The first branch includes a first convolutional upsampling block, a second convolutional upsampling block, a third convolutional upsampling block, and a fourth convolutional upsampling block connected in sequence. The second branch includes a first bidirectional cross-attention module, a first convolutional upsampling block, a second bidirectional cross-attention module, a second convolutional upsampling block, a third bidirectional cross-attention module, a third convolutional upsampling block, a fourth bidirectional cross-attention module, and a fourth convolutional upsampling block connected in sequence. The first convolutional upsampling block in the first branch receives the feature map and generates a feature map for output. The second convolutional upsampling block receives the feature map and generates a feature map for output. The third convolutional upsampling block receives the feature map and generates a feature map for output. The first bidirectional cross-attention module in the second branch receives the feature map and the feature map and generates a feature map through the first convolutional upsampling block. The second bidirectional cross-attention module receives the feature map and the feature map and generates a feature map through the second convolutional upsampling block. The third bidirectional cross-attention module receives the feature map and the feature map and generates a feature map through the third convolutional upsampling block. The fourth bidirectional cross-attention module receives the feature map and the feature map and generates the reconstructed image through the fourth convolutional upsampling block. The size of the feature map is , the size of the feature map is , the size of the feature map is , the size of the feature map is ; the size of the feature map is , the size of the feature map is , the size of the feature map is , the size of the feature map is .
Citation Information
Patent Citations
Hyperspectral fusion imaging method and system based on internal and external priori, and medium
CN115311187A
Image super-resolution reconstruction method and related equipment
CN119205507A
Underwater image enhancement quality evaluation method based on quality perception domain adaptation
CN119313661A
An edge-guided RGBD underwater salient object detection method with multi-attention
JP7605548B1
Device, system, and method of blind deblurring and blind super-resolution utilizing internal patch recurrence
US20140354886A1