Reversible network-based lightweight image just noticeable distortion prediction method
Through a lightweight image perceptible distortion prediction method based on reversible network, a pseudo-label is generated using a high dynamic range visual difference predictor, combining feature enhancement and discrete wavelet transformation, the high-cost annotation problem is solved, and high-quality pixel-level perceptible distortion threshold prediction is achieved, which improves the robustness and generalization ability of the model.
Patent Information
- Application Number
- CN202510772154.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-11
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2045-06-11
AI Technical Summary
In the prior art, distortion threshold labeling is easily perceived with high cost and lack of pixel-level prediction, resulting in a decrease in practicality of image-level prediction results.
A lightweight image perceptible distortion prediction method based on a reversible network is adopted to generate pseudo-labels through a high dynamic range visual difference predictor. Combining feature enhancement, discrete wavelet transformation and high-low frequency image processing, an image perceptible distortion threshold prediction network is built to achieve pixel-level prediction.
It significantly reduces the labeling cost, provides high-quality pixel-level, accurate distortion threshold prediction, and improves the robustness and generalization capabilities of the prediction model.
Smart Images

Figure CN120299043A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image distortion prediction. Specifically, it relates to a lightweight image just noticeable distortion prediction method based on a reversible network. Background Art
[0002] In the fields of psychology and perceptual science, just noticeable difference (JND) is an important concept used to describe the smallest stimulus change that the human visual system can perceive. This concept was first proposed by the German psychologist Gustav Fechner in his psychophysics research and has become one of the bases for understanding human perceptual abilities.
[0003] By studying just noticeable distortion, it can help us better understand how humans perceive changes in visual information, so as to optimize applications such as image compression, image enhancement, and information hiding. In recent years, deep learning technology has made breakthrough progress, and models represented by convolutional neural networks (CNNs) have demonstrated excellent performance in the field of computer vision.
[0004] The prediction of just noticeable distortion thresholds now faces two major challenges: Manually annotating pixel-level just noticeable distortion thresholds has problems of strong subjectivity and high cost, and there is a lack of a suitable end-to-end training dataset. In response to this technical bottleneck, some recent studies have adopted a compromise strategy, modeling the just noticeable distortion threshold prediction as a binary classification task. However, these image-level just noticeable distortion threshold predictors cannot provide pixel-level just noticeable distortion thresholds, resulting in a reduction in the practicality of the prediction results. Pixel-level just noticeable distortion thresholds are essentially necessary in various perceptual image processing applications. Therefore, there is an urgent need to study a method that can reduce the annotation cost and achieve lightweight pixel-level just noticeable distortion prediction. Summary of the Invention
[0005] The technical problem to be solved by the present invention is how to significantly reduce the just noticeable distortion threshold annotation cost and achieve lightweight pixel-level just noticeable distortion threshold prediction. To overcome the defects of the above prior art (or related art), the present invention provides a lightweight image just noticeable distortion prediction method based on a reversible network.
[0006] The present invention provides a lightweight image just noticeable distortion prediction method based on a reversible network, including the following steps: Step S1, obtain multiple original images and input them into a high-dynamic-range visual difference predictor to construct a training set and a test set, and preprocess the training set to obtain multiple Y-channel images; Step S2: Build an image just-noticeable distortion threshold prediction network using a deep learning framework. The low-frequency feature map is obtained by sequentially performing feature enhancement, discrete wavelet transform, and high- and low-frequency image extraction on the Y-channel image through the image just-noticeable distortion threshold prediction network. and the high-frequency feature map . And perform high-frequency modulation on the high-frequency feature map to obtain a high-frequency feature map . Also, perform high- and low-frequency image inversion, discrete wavelet transform, and feature enhancement on the low-frequency feature map and the high-frequency feature map in sequence to obtain a critically perceptually lossless image; Step S3: Train the image just-noticeable distortion threshold prediction network using the training set according to Step S2 to obtain an image just-noticeable distortion threshold prediction model; Step S4: Test the test set through the image just-noticeable distortion threshold prediction model to obtain the critically perceptually lossless image corresponding to the test set, and obtain an absolute difference map as the predicted just-noticeable distortion threshold map based on the pairwise corresponding critically perceptually lossless image and Y-channel image of the test set.
[0007] Compared with the prior art, a lightweight image just-noticeable distortion prediction method based on a reversible network of the present invention has the following advantages: In the present invention, the just-noticeable distortion threshold is redefined as the absolute difference between the original Y-channel image and its critically perceptually lossless image, providing a new idea for the prediction of pixel-level just-noticeable distortion thresholds. The high-dynamic-range visual difference predictor is used to predict and obtain pseudo-label partitions to generate the training set and the test set, effectively solving the problem of high annotation costs faced when using large-scale original images as true labels. By building an image just-noticeable distortion threshold prediction network and focusing on the modulation of high-frequency information, based on the reversible characteristics of the image just-noticeable distortion threshold prediction network, through a forward processing process including feature enhancement, discrete wavelet transform, and high- and low-frequency image extraction and a reverse processing process including high- and low-frequency image inversion, discrete wavelet transform, and feature enhancement, the corrected high-frequency information is fused with the intact low-frequency information to obtain a critically perceptually lossless image, realizing high-quality image prediction.
[0008] In a possible implementation manner, in Step S1, each of the original images is respectively subjected to Karhunen - Loève transform to generate degraded version images with different degrees, and the perceptual distortion probability between each original image and the corresponding degraded version image is predicted using the high-dynamic-range visual difference predictor. The degraded version images with a perceptual distortion probability of 0.75 are selected to divide the training set and the test set.
[0009] In a possible implementation manner, in the step S1, the process of preprocessing the training set includes: The degraded version images in the training set are uniformly cropped with a fixed size of 256×256 pixels to obtain cropped images, and then each cropped image is converted into the YUV color space and the Y-channel luminance component therein is extracted to obtain the Y-channel image.
[0010] In a possible implementation manner, the just-noticeable image distortion threshold prediction network constructed in the step S2 includes a feature enhancement module, a discrete wavelet transform module, a reversible module, and a high-frequency modulation module connected in sequence. The feature enhancement module receives the Y-channel image to perform feature enhancement to obtain a feature map , and the discrete wavelet transform module decomposes the feature map into a low-frequency feature map and a high-frequency feature map . The reversible module performs high-low frequency image extraction on the low-frequency feature map and the high-frequency feature map to obtain the low-frequency feature map and the high-frequency feature map . The high-frequency modulation module modulates the high-frequency feature map to obtain the high-frequency feature map . The reversible module performs high-low frequency image inversion on the low-frequency feature map and the high-frequency feature map to obtain the low-frequency feature map and the high-frequency feature map . The discrete wavelet transform module performs an inverse discrete wavelet transform on the low-frequency feature map and the high-frequency feature map to obtain a reconstructed feature map . The feature enhancement module performs feature enhancement on the feature map to obtain the critically perceptually lossless image.
[0011] In a possible implementation manner, the feature enhancement module includes a first dense block, a first convolutional layer, and a first splicing layer connected in sequence. The first dense block includes a first convolutional layer, a second convolutional layer, a third convolutional layer, and a fourth convolutional layer connected in sequence. The first convolutional layer receives the Y-channel image to perform a convolutional operation to obtain a feature map and adds the feature map and the Y-channel image in channels to obtain a feature map . The second convolutional layer receives the feature map Perform a convolution operation to obtain a feature map And use the said feature map And the said feature map Perform channel addition to obtain a feature map , Receive the said feature map through the third convolutional layer Perform a convolution operation to obtain a feature map And use the said feature map And the said feature map Perform channel addition to obtain a feature map , Receive the said feature map through the fourth convolutional layer Perform a convolution operation to obtain a feature map , Through the fourth convolutional layer, use the said feature map And the said feature map Perform channel addition and output channel setting to 1 through the first convolutional layer to obtain a feature map , Through the first splicing layer, perform residual connection on the said feature map And the Y-channel image to obtain the said feature map .
[0012] In a possible implementation, the feature enhancement module in step S2 includes a second dense block, a second convolutional layer, and a second splicing layer connected in sequence. The second dense block includes a first convolutional layer, a second convolutional layer, a third convolutional layer, and a fourth convolutional layer connected in sequence. Receive the said feature map through the first convolutional layer Perform a convolution operation to obtain a feature map And use the said feature map And the said feature map Perform channel addition to obtain a feature map , Receive the said feature map through the second convolutional layer Perform a convolution operation to obtain a feature map And use the said feature map And the said feature map Perform channel addition to obtain a feature map , Receive the said feature map through the third convolutional layer Perform a convolution operation to obtain a feature map And use the said feature map And the said feature map Perform channel addition to obtain a feature map , Receive the said feature map through the fourth convolutional layer Perform a convolution operation to obtain a feature map , Through the fourth convolutional layer, use the said feature map And the said feature map Perform channel addition and set the output channels to 1 through the second convolutional layer to obtain a feature map , and through the second splicing layer, splice the feature map and the feature map for residual connection to obtain the critically perceived lossless image.
[0013] In a possible implementation, the reversible module includes four reversible blocks connected in sequence. Through the four reversible blocks, the low-frequency feature map and the high-frequency feature map are successively subjected to high-low frequency image extraction to obtain the low-frequency feature map and the high-frequency feature map ; or Through the four reversible blocks, the low-frequency feature map and the high-frequency feature map are successively subjected to high-low frequency image inversion to obtain the low-frequency feature map and the high-frequency feature map .
[0014] In a possible implementation, the high-frequency modulation module includes a first convolutional layer, a second convolutional layer, and a third convolutional layer connected in sequence. Through the first convolutional layer, the high-frequency feature map is first modulated to obtain a high-frequency feature map , through the second convolutional layer, the high-frequency feature map is secondarily modulated to obtain a high-frequency feature map , and through the third convolutional layer, the high-frequency feature map is tertially modulated to obtain the high-frequency feature map .
[0015] In a possible implementation, the convolutional kernel size of the first convolutional layer of the high-frequency modulation module is 1, the padding size is 0, the number of input channels is 3, and the number of output channels is 128; the convolutional kernel size of the second convolutional layer is 3, the padding size is 1, the number of input channels is 128, and the number of output channels is 128; the convolutional kernel size of the third convolutional layer is 1, the padding size is 0, the number of input channels is 128, and the number of output channels is 3. Description of the Drawings
[0016] Figure 1 is the flowchart of the steps of the present invention; Figure 2 is the framework schematic diagram of the image just noticeable distortion threshold prediction network of the present invention; Figure 3 is the framework schematic diagram of the dense block in the feature enhancement module of the present invention; Figure 4 Schematic diagram of the prediction effects of existing just-noticeable distortion threshold prediction methods Yang05, Wu13, Wu17, and Jakh18; Figure 5 Schematic diagram of the prediction effects of the method of the present invention and existing just-noticeable distortion threshold prediction methods Chen19, Wang21, and Jiang22. Detailed implementation manners
[0017] First of all, those skilled in the art should understand that these implementation manners are only used to explain the technical principles of the embodiments of the present invention, and are not intended to limit the protection scope of the embodiments of the present invention. Those skilled in the art can make adjustments according to needs to adapt to specific application scenarios.
[0018] The present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments.
[0019] See Figure 1 , the embodiment of the present invention discloses a lightweight image just-noticeable distortion prediction method based on a reversible network, including the following steps: Step S1, obtaining multiple original images and inputting them into a high-dynamic-range visual difference predictor to construct a training set and a test set, and preprocessing the training set to obtain multiple Y-channel images; Step S2, using a deep learning framework to build an image just-noticeable distortion threshold prediction network, and sequentially performing feature enhancement, discrete wavelet transform, and high-frequency and low-frequency image extraction on the Y-channel images through the image just-noticeable distortion threshold prediction network to obtain a low-frequency feature map and a high-frequency feature map , and performing high-frequency modulation on the high-frequency feature map to obtain a high-frequency feature map , and sequentially performing high-frequency and low-frequency image inversion, discrete wavelet transform, and feature enhancement on the low-frequency feature map and the high-frequency feature map to obtain a critically-perceptual lossless image; Step S3, training the image just-noticeable distortion threshold prediction network using the training set according to Step S2 to obtain an image just-noticeable distortion threshold prediction model; Step S4, testing the test set through the image just-noticeable distortion threshold prediction model to obtain the critically-perceptual lossless image corresponding to the test set, and obtaining an absolute difference map as the predicted just-noticeable distortion threshold map according to the pairwise corresponding critically-perceptual lossless image and Y-channel image of the test set.
[0020] In the embodiment of the present application, in step S1, a pixel-level just noticeable distortion threshold dataset is constructed. To solve the problems of strong subjectivity and high cost in manually annotating the pixel-level just noticeable distortion threshold, a high-dynamic-range visual difference predictor that can predict the perceptual distortion probability based on an image pair containing the original image and the distorted image is introduced to provide pseudo-labels as the training set and the validation set. At the same time, a part of the original images are manually annotated as the test set.
[0021] In the embodiment of the present application, aiming at the two core defects of the traditional manual annotation method: the annotation inconsistency caused by subjective perceptual differences and the high time cost required for pixel-level annotation, a high-dynamic-range visual difference predictor is innovatively introduced. By simulating the multi-channel perception mechanism of the human visual system and combining the adaptive contrast sensitivity function and the luminance perception non-linear compensation algorithm, the high-dynamic-range visual difference predictor can perform pixel-level visual difference probability prediction (range 0-1) on the input original image-distorted image pair, generating pseudo JND annotation data with perceptual consistency. In the specific implementation process, each original image undergoes Karhunen-Loeve transform (KLT transform) to generate 256 different degrees of degraded version images, and the high-dynamic-range visual difference predictor is used to sequentially predict the perceptual distortion probability between the original image and its degraded version. The perceptual distortion probability of 0.75 (that is, the probability of being predicted as distorted is 75%) is selected as the critical perceptual lossless image corresponding to the original image. Subsequently, the high-dynamic-range visual difference predictor is used to batch process 8965 groups of critical perceptual lossless images to construct a large-scale pseudo-labeled training set (70%) and a validation set (10%), and at the same time, 20% of the manually annotated real data is reserved as the test set to ensure that the image just noticeable distortion threshold prediction model has reliable real-scene generalization ability while maintaining the advantages of automation. This hybrid annotation strategy not only significantly improves the data preparation efficiency but also enhances the robustness of the image just noticeable distortion threshold prediction model.
[0022] In the embodiment of the present application, since the input of the deep neural network usually has a fixed form, it is necessary to preprocess the critical perceptual lossless images in the test set. Specifically, the critical perceptual lossless images are uniformly cropped to , to ensure that the images input into the deep neural network have consistent dimensions. Subsequently, the critical perceptual lossless images in RGB format are converted to the YUV color space, and the luminance component (Y channel) is extracted as the input feature of the neural network. This processing method not only reduces the computational complexity but also retains the main structural information of the critical perceptual lossless images, which is beneficial to subsequent feature learning and model training.
[0023] In the embodiment of the present application, in step S2, a deep learning framework is used to build an image just noticeable distortion threshold prediction network based on a reversible network, such asFigure 2 As shown, the just-noticeable distortion threshold prediction network for the image includes a feature enhancement module, a discrete wavelet transform module, a reversible module, and a high-frequency modulation module, and has two processes: forward processing and reverse processing.
[0024] In the embodiment of the present application, the feature enhancement module is distributed at the entrance and exit of the just-noticeable distortion threshold prediction network for the image. In the forward processing process, the input image is a Y-channel image with a size of. The feature enhancement module includes 1 first dense block and 1 first convolutional layer. There are 4 convolutional layers in the first dense block, and 16 channels are added in each layer. The output channel number of the previous layer is added to the input channel number of the current layer. Here, the feature addition operation is a conventional operation in the convolutional neural network; let the Y-channel image be , and the feature map output by the first convolutional layer in the first dense block is . The feature maps input and output by the second convolutional layer are and respectively. The feature maps input and output by the third convolutional layer are and respectively. The feature maps input and output by the fourth convolutional layer are and respectively. After the first dense block is processed, the number of channels is adjusted to 1 (the same as the output channel number of the residual connection) through the first convolutional layer. The output channel number of this convolutional operation is 1, and the feature map is obtained. Finally, in the first splicing layer, feature addition is performed with the identity mapping of the Y-channel image . The Y-channel image and the feature map output after feature enhancement are added to obtain the feature map ; since the input and output channel numbers match, the residual connection does not change the size. Specifically, as Figure 3 shown, where the size of the Y-channel image is , the size of the feature map is , the size of the feature map is , the size of the feature map is , the size of the feature map is , the size of the feature map is , the size of the feature map is , the size of the feature map is , the size of the feature map is , the feature map has a size of , in this embodiment .
[0025] In the embodiment of the present application, the complete reversibility of the discrete wavelet transform module is an important feature, ensuring the lossless transmission of information during the forward and inverse transformation processes of the feature map. This feature makes it of great value in many image processing tasks. In the embodiment of the present invention, the discrete wavelet transform module can decompose the feature map into a low-frequency feature map and a high-frequency feature map , where the low-frequency feature map has a size of , and the high-frequency feature map has a size of .
[0026] In the embodiment of the present application, the input of the reversible module is the low-frequency feature map and the high-frequency feature map output by the discrete wavelet transform module, and the output is the low-frequency feature map and the high-frequency feature map . The reversible module contains 4 reversible blocks, each reversible block contains multiple convolutional blocks, and each convolutional block is composed of a convolutional layer consisting of two convolutions and the activation function ReLU. As Figure 2 shown, taking the first reversible block in the forward processing process as an example, its forward processing process is as follows: For the input high-frequency feature map pass through the convolutional block to obtain the feature map , pass through the convolutional block to obtain the feature map , then the feature map is multiplied by the low-frequency feature map to obtain the feature map . Then, the feature map is added to the feature map to obtain the low-frequency feature map output by the first reversible block; to prevent the occurrence of gradient disappearance or gradient explosion during the training process, the low-frequency feature map is subjected to a Sigmoid activation operation to obtain the high-frequency feature map ; for the input high-frequency feature map pass through the convolutional block to obtain the feature map , pass through the convolutional block to obtain the feature map , then the feature map is multiplied by the high-frequency feature map to obtain the feature map Subsequently, the feature map is added to the feature map to obtain the high-frequency feature map of the output of the first reversible block ; Subsequently, the low-frequency feature map and the high-frequency feature map are input into the subsequent three reversible blocks with the same structure but different parameters, that is, a total of 4 reversible block operations are performed, and the low-frequency feature map and the high-frequency feature map are output, where the feature map , the feature map , the feature map , the low-frequency feature map , the high-frequency feature map , the low-frequency feature map have a size of , and the feature map , the feature map , the feature map , the high-frequency feature map , the high-frequency feature map have a size of .
[0027] In the embodiment of the present application, the high-frequency modulation module is composed of three cascaded convolutional layers, which are responsible for modulating the high-frequency feature map output by the reversible module, aiming to potentially learn the relationship between the just noticeable distortion threshold and the high-frequency information. Specifically, the high-frequency feature map is input into the first convolutional layer to obtain the high-frequency feature map , input into the second convolutional layer to obtain the high-frequency feature map , and finally input into the third convolutional layer to obtain the modulated high-frequency feature map , where the high-frequency feature map has a size of , the high-frequency feature map has a size of , and the high-frequency feature map has a size of .
[0028] In the embodiment of the present application, the high-frequency modulation module is composed of three cascaded convolutional layers with convolution kernels of 1, 3, and 1. Specifically, the first convolutional layer has a convolution kernel size of 1, a padding size of 0, an input channel number of 3, and an output channel number of 128; the second convolutional layer has a convolution kernel size of 3, a padding size of 1, an input channel number of 128, and an output channel number of 128; the third convolutional layer has a convolution kernel size of 1, a padding size of 0, an input channel number of 128, and an output channel number of 3.
[0029] In the embodiments of the present application, after completing all the forward propagation processes, the modulated high-frequency feature map and the low-frequency feature map are used as the input of the reverse processing process of the reversible module. Taking the first reversible block of the reverse processing process as an example, its reverse processing process is as follows: For the input low-frequency feature map pass through the convolutional block to obtain the feature map , pass through the convolutional block to obtain the feature map . Subsequently, the low-frequency feature map subtracts the feature map to obtain the feature map , and then divides by the feature map to obtain the high-frequency feature map ; to prevent the occurrence of gradient disappearance or gradient explosion during the training process, perform a Sigmoid activation operation on the low-frequency feature map to obtain the high-frequency feature map ; for the input high-frequency feature map pass through the convolutional block to obtain the feature map , pass through the convolutional block to obtain the feature map . Subsequently, the high-frequency feature map subtracts the feature map to obtain the feature map , and then divides by the feature map to obtain the low-frequency feature map ; Similarly, the low-frequency feature map and the high-frequency feature map are used as the input of the next reversible block and input into the subsequent three reversible blocks with the same structure but different parameters, that is, a total of 4 reversible block inverse operations are performed, and the output is the low-frequency feature map and the high-frequency feature map . Among them, the size of the feature map , the feature map , the feature map , and the low-frequency feature map is , and the size of the feature map , the feature map , the feature map , the high-frequency feature map , and the high-frequency feature map is ; subsequently, the low-frequency feature map and the high-frequency feature map pass through the discrete wavelet transform module for inverse transformation to obtain the reconstructed feature map , finally, through the feature enhancement module at the output, the predicted critical perception lossless image is obtained , where the feature map and the critical perception lossless image have dimensions of .
[0030] In the embodiment of the present application, during the reverse processing, the input image of the feature enhancement module is a feature map with dimensions of . The feature enhancement module includes 1 second dense block and 1 second convolutional layer. There are 4 convolutional layers in the second dense block, and 16 channels are added in each layer, and the number of output channels of the previous layer is added to the number of input channels of the current layer. Let the feature map output by the first convolutional layer in the second dense block be , and the feature maps input and output by the second convolutional layer are respectively and , the feature maps input and output by the third convolutional layer are respectively and , the feature maps input and output by the fourth convolutional layer are respectively and . After the second dense block is processed, the number of channels is adjusted to 1 (the same as the number of output channels of the residual connection) through the second convolutional layer. The output channel number of this convolutional operation is 1, and the feature map is obtained. Finally, in the second splicing layer, by performing feature addition with the identity mapping of the feature map , the feature map and the feature map output after feature enhancement are added to obtain the critical perception lossless image ; since the input and output channel numbers match, the residual connection does not change the dimensions. Among them, the dimension of the feature map is , the dimension of the feature map is , the dimension of the feature map is , the dimension of the feature map is , the dimension of the feature map is , the dimension of the feature map is , the dimension of the feature map is , the dimension of the feature map is , in this embodiment .
[0031] In the embodiment of the present application, in step S3, a training set is used to train the image just-noticeable distortion threshold prediction network based on the reversible network. After each round of training, according to the output critically-perceptual lossless image and the input Y-channel image the loss of the network is calculated, denoted as , , where MSE represents the mean squared error loss. The initial learning rate of the SGD optimizer is 0.01, and the network parameters are adjusted using a batch size of 16.
[0032] In the embodiment of the present application, in step S3, the image just-noticeable distortion threshold prediction network is trained for a total of 200 rounds according to the process of step S2. At the 100th round, the learning rate of the SGD optimizer decays to 0.001, and finally, an image just-noticeable distortion threshold prediction network model based on the reversible network is obtained through training.
[0033] In the embodiment of the present application, both the first dense block and the second dense block in the feature enhancement module have 4 convolutional blocks. Each layer adds 16 channels, and the number of output channels of the previous layer is added to the number of input channels of the current layer. Finally, a convolutional block restores the number of channels to 1. Specifically, the convolutional kernel size of the first convolutional block is 3, the padding size is 1, the number of input channels is 1, and the number of output channels is 16. The number of output channels of the first layer is added to the number of input channels of the current layer to obtain the number of input channels of the second layer as 17; the convolutional kernel size of the second convolutional block is 3, the padding size is 1, the number of input channels is 17, and the number of output channels is 16. The number of output channels of the second layer is added to the number of input channels of the current layer to obtain the number of input channels of the third layer as 33; the convolutional kernel size of the third convolutional block is 3, the padding size is 1, the number of input channels is 33, and the number of output channels is 16. The number of output channels of the third layer is added to the number of input channels of the current layer to obtain the number of input channels of the fourth layer as 49; the convolutional kernel size of the fourth convolutional block is 3, the padding size is 1, the number of input channels is 49, and the number of output channels is 16. The convolutional kernel size of the last convolutional block is 1, the padding size is 0, the number of input channels is 16, and the number of output channels is 1. Finally, the output features of the last convolutional block are added to the input features to obtain the final output.
[0034] In the embodiment of the present application, the discrete wavelet transform module can decompose the input feature map into low-frequency and high-frequency signal components. For a 2D image, the discrete wavelet transform is mainly decomposed into the following four sub-bands: the low-frequency part, the horizontal high-frequency part, the vertical high-frequency part, and the diagonal high-frequency part. In the embodiment of the present invention, the haar wavelet basis is used to perform wavelet transformation on the feature map output by the feature enhancement module to obtain the low-frequency part, the horizontal high-frequency part, the vertical high-frequency part, and the diagonal high-frequency part of the feature map, and then a low-frequency feature map and a high-frequency feature map are generated.
[0035] In the embodiments of the present application, 4 reversible blocks are used in the reversible module, and each reversible block contains 4 convolutional blocks. Since these four reversible blocks have the same structure but different weights, only the specific structure of one reversible block will be detailed below. The convolutional blocks mentioned above , convolutional block , convolutional block , convolutional block all contain two convolutional layers, and each convolutional layer is connected to a ReLU activation function. Specifically, the convolutional kernel size of the first convolutional layer in convolutional block is 1, the padding size is 0, the number of input channels is 3, and the number of output channels is 64; the convolutional kernel size of the second convolutional layer in convolutional block is 1, the padding size is 0, the number of input channels is 64, and the number of output channels is 1; the convolutional kernel size of the first convolutional layer in convolutional block is 1, the padding size is 0, the number of input channels is 3, and the number of output channels is 64; the convolutional kernel size of the second convolutional layer in convolutional block is 1, the padding size is 0, the number of input channels is 64, and the number of output channels is 1; the convolutional kernel size of the first convolutional layer in convolutional block is 1, the padding size is 0, the number of input channels is 1, and the number of output channels is 64; the convolutional kernel size of the second convolutional layer in convolutional block is 1, the padding size is 0, the number of input channels is 64, and the number of output channels is 3; the convolutional kernel size of the second convolutional layer in convolutional block is 1, the padding size is 0, the number of input channels is 1, and the number of output channels is 64; the convolutional kernel size of the second convolutional layer in convolutional block is 1, the padding size is 0, the number of input channels is 64, and the number of output channels is 3.
[0036] To further verify the feasibility and effectiveness of the method of the present invention, the following experiments are conducted on the method of the present invention: During the experiment, the pixel-level just-noticeable distortion threshold dataset constructed by the method of the present invention is directly selected. To further verify the generalization ability of the image just-noticeable distortion threshold prediction model, this experiment is further tested on the existing visually lossless threshold (VLT) dataset (Mikhailiuk et al., 2021); First, the performance of the image just-noticeable distortion threshold prediction model proposed by the method of the present invention in just-noticeable distortion threshold prediction is evaluated. Table 1 below summarizes the experimental results of this image just-noticeable distortion threshold prediction model on the test set of the just-noticeable distortion threshold dataset proposed by the method of the present invention: Table 1 Root mean square error result index table of the image just-noticeable distortion threshold prediction model in the method of the present invention ; Then, it is classified according to the scene type. Among them, there are several key observations worthy of attention. It should be noted that the method of the present invention has achieved the lowest root mean square error in all scene categories and the entire test set, demonstrating state-of-the-art performance. This verifies the effectiveness of constructing critical perception lossless images with pseudo-labels for training and emphasizes the importance of paying attention to high-frequency information in the prediction of just noticeable distortion thresholds; To further evaluate the generalization ability of the just noticeable distortion threshold prediction model for images, its performance was tested on the VLT dataset. The experimental results are shown in Table 2 below: Table 2 Root Mean Square Error Result Index Table of the Method of the Present Invention on the VLT Dataset ; As can be seen from Table 2 above, in this experiment, the just noticeable distortion threshold prediction model proposed by the method of the present invention was only trained on the pseudo-labels proposed by the method of the present invention and then directly tested on the VLT dataset without any fine-tuning. Although the just noticeable distortion threshold prediction model is completely trained based on pseudo-label data, 5 out of 20 test images outperformed all competing methods, and it achieved the second-best result in terms of overall average performance. These results indicate that the just noticeable distortion threshold prediction model proposed by the method of the present invention has strong generalization ability and can effectively adapt to unknown data from different image datasets; In addition to directly comparing the prediction performance, the effectiveness of the just noticeable distortion threshold prediction model for images can also be indirectly verified by evaluating its noise hiding ability. For a given image, first use the just noticeable distortion threshold prediction model for images to calculate its corresponding visible just noticeable distortion threshold. Then, use the calculated just noticeable distortion threshold as the adjustment factor for noise embedding, inject Gaussian white noise with the same intensity (peak signal-to-noise ratio equal to 26 dB) into the Y-channel image, and conduct a visual quality comparative analysis. Figure 4 and Figure 5 shows the schematic diagrams of the prediction effects corresponding to the method of the present invention and the existing just noticeable distortion threshold prediction methods (Yang05, Wu13, Wu17, Jakh18, Chen19, Wang21, Jiang22), as well as the noise injection images guided by each method. The results show that the just noticeable distortion threshold prediction model proposed by the method of the present invention is superior to other comparison methods and exhibits more excellent performance.
[0037] In the description of the present invention, the descriptions referring to terms such as "one embodiment", "some embodiments", "in this embodiment", "specific examples", or "some examples", etc. mean that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic expressions of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in a suitable manner in any one or more embodiments or examples. In addition, without contradiction, those skilled in the art can combine and combine the different embodiments or examples described in this specification and the features of different embodiments or examples.
[0038] As described above, it is only the specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any changes or substitutions that can be easily thought of by those skilled in the art within the technical scope disclosed by the present invention should be covered by the protection scope of the present invention. Therefore, the protection scope of the present invention should be subject to the protection scope of the claims.
Claims
1. A lightweight image just noticeable distortion prediction method based on a reversible network, characterized in that Including the following steps: Step S1: Obtain multiple original images and input them into a high-dynamic-range visual difference predictor to construct a training set and a test set, and preprocess the training set to obtain multiple Y-channel images; Step S2: Build an image just noticeable distortion threshold prediction network using a deep learning framework, and sequentially perform feature enhancement, discrete wavelet transform, and high-frequency and low-frequency image extraction on the Y-channel image through the image just noticeable distortion threshold prediction network to obtain a low-frequency feature map and a high-frequency feature map , and perform high-frequency modulation on the high-frequency feature map to obtain a high-frequency feature map , and sequentially perform high-frequency and low-frequency image inversion, discrete wavelet transform, and feature enhancement on the low-frequency feature map and the high-frequency feature map to obtain a critically perceptually lossless image; Step S3: Train the just-noticeable distortion threshold prediction network for images using the training set according to Step S2 to obtain a just-noticeable distortion threshold prediction model for images; Step S4: Test the test set through the just-noticeable distortion threshold prediction model for images to obtain the corresponding critically perceptually lossless images of the test set, and obtain an absolute difference map as the predicted just-noticeable distortion threshold map based on the pairwise corresponding critically perceptually lossless images and Y-channel images of the test set.
2. The lightweight image just noticeable distortion prediction method based on a reversible network according to claim 1, wherein, In Step S1, perform Karhunen-Loève transform on each of the original images to generate degraded version images with different degrees, and use the high-dynamic-range visual difference predictor to predict the perceptual distortion probability between each of the original images and the corresponding degraded version images, and select the degraded version images with a perceptual distortion probability of 0.75 to divide the training set and the test set.
3. The method for predicting the just noticeable distortion of a lightweight image based on a reversible network according to claim 2, wherein In Step S1, the process of preprocessing the training set includes: Uniformly crop the degraded version images in the training set using a fixed size of 256×256 pixels to obtain cropped images, and then convert each of the cropped images to the YUV color space and extract the Y-channel luminance component therein to obtain the Y-channel images.
4. The lightweight image just noticeable distortion prediction method based on a reversible network according to claim 1, characterized in that, The just-noticeable distortion threshold prediction network constructed in the step S2 includes a feature enhancement module, a discrete wavelet transform module, a reversible module, and a high-frequency modulation module connected in sequence. The Y-channel image is received by the feature enhancement module for feature enhancement to obtain a feature map , and the feature map is decomposed into a low-frequency feature map and a high-frequency feature map by the discrete wavelet transform module. The low-frequency feature map and the high-frequency feature map are used by the reversible module to extract high and low frequency images to obtain the low-frequency feature map and the high-frequency feature map . The high-frequency feature map is modulated by the high-frequency modulation module to obtain the high-frequency feature map . The low-frequency feature map and the high-frequency feature map are used by the reversible module to perform high and low frequency image inversion to obtain the low-frequency feature map and the high-frequency feature map . The discrete wavelet transform inverse transform is performed on the low-frequency feature map and the high-frequency feature map by the discrete wavelet transform module to obtain a reconstructed feature map . The feature map is enhanced by the feature enhancement module to obtain the critically perceptually lossless image.
5. The method for predicting the just noticeable distortion of a lightweight image based on a reversible network according to claim 4, wherein The feature enhancement module includes a first dense block, a first convolutional layer, and a first splicing layer connected in sequence. The first dense block includes a first convolutional layer, a second convolutional layer, a third convolutional layer, and a fourth convolutional layer connected in sequence. The first convolutional layer receives the Y-channel image for convolution operation to obtain a feature map and adds the feature map and the Y-channel image in channels to obtain a feature map . The second convolutional layer receives the feature map for convolution operation to obtain a feature map and adds the feature map and the feature map in channels to obtain a feature map . The third convolutional layer receives the feature map for convolution operation to obtain a feature map and adds the feature map and the feature map in channels to obtain a feature map . The fourth convolutional layer receives the feature map for convolution operation to obtain a feature map . The fourth convolutional layer adds the feature map and the feature map in channels and outputs the channel as 1 through the first convolutional layer to obtain a feature map . The first splicing layer performs residual connection on the feature map and the Y-channel image to obtain the feature map .
6. The lightweight image just noticeable distortion prediction method based on a reversible network according to claim 4, wherein The feature enhancement module in the step S2 includes a second dense block, a second convolutional layer, and a second splicing layer connected in sequence. The second dense block includes a first convolutional layer, a second convolutional layer, a third convolutional layer, and a fourth convolutional layer connected in sequence. The feature map is received through the first convolutional layer to perform a convolution operation to obtain a feature map and the feature map and the feature map are added channel by channel to obtain a feature map , and the feature map is received through the second convolutional layer to perform a convolution operation to obtain a feature map and the feature map and the feature map are added channel by channel to obtain a feature map , and the feature map is received through the third convolutional layer to perform a convolution operation to obtain a feature map and the feature map and the feature map are added channel by channel to obtain a feature map , and the feature map is received through the fourth convolutional layer to perform a convolution operation to obtain a feature map , through the fourth convolutional layer, the feature map and the feature map are added channel by channel and the output channels are set to 1 through the second convolutional layer to obtain a feature map , through the second splicing layer, the feature map and the feature map are subjected to residual connection to obtain the critical perception lossless image.
7. The method for predicting the just noticeable distortion of a lightweight image based on a reversible network according to claim 4, characterized in that The reversible module includes four reversible blocks connected in sequence, and the low-frequency feature map and the high-frequency feature map are successively subjected to high-low frequency image extraction to obtain the low-frequency feature map and the high-frequency feature map ; or Through the four reversible blocks, the low-frequency feature map and the high-frequency feature map are successively subjected to high-low frequency image inversion to obtain the low-frequency feature map and the high-frequency feature map .
8. The method for predicting lightweight image just noticeable distortion based on a reversible network according to claim 4, characterized in that The high-frequency modulation module includes a first convolutional layer, a second convolutional layer, and a third convolutional layer connected in sequence. The first convolutional layer performs a first modulation on the high-frequency feature map to obtain a high-frequency feature map . The second convolutional layer performs a second modulation on the high-frequency feature map to obtain a high-frequency feature map . The third convolutional layer performs a third modulation on the high-frequency feature map to obtain the high-frequency feature map .
9. The lightweight image just noticeable distortion prediction method based on a reversible network according to claim 8, wherein The convolution kernel size of the first convolution layer of the high-frequency modulation module is 1, the padding size is 0, the number of input channels is 3, and the number of output channels is 128; the convolution kernel size of the second convolution layer of the high-frequency modulation module is 3, the padding size is 1, the number of input channels is 128, and the number of output channels is 128; the convolution kernel size of the third convolution layer of the high-frequency modulation module is 1, the padding size is 0, the number of input channels is 128, and the number of output channels is 3.
Citation Information
Patent Citations
Top-down natural image just noticeable distortion threshold estimation method
CN114519668A
Non-reference image quality evaluation method based on multi-scale feature fusion
CN116128807A
Multi-level full RGB just noticeable perceptual coding distortion prediction method for video compression
CN119135909A
Three-dimensional point cloud just noticeable distortion threshold prediction method and product
CN119919340A