Image quality evaluation method based on data enhancement and contrast learning

By using data augmentation and contrast learning techniques in image quality evaluation, an image quality evaluation model based on U-Net encoder was constructed, which solved the problems of large sample requirements, large model parameter scale and high training difficulty in the prior art, and achieved higher evaluation accuracy and interpretability.

CN119941691AActive Publication Date: 2025-05-06WUHAN UNIV

Patent Information

Application Number
CN202510060471.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-15
Publication Date
2025-05-06
Estimated Expiration
2045-01-15

AI Technical Summary

Technical Problem

The existing deep learning image quality evaluation algorithm requires large samples, large scale of model parameters, high training difficulty, and lack interpretability.

Method used

The image quality evaluation method based on data augmentation and contrast learning is adopted, and preprocessed through random data augmentation, deep image priors and bilinear interpolation, combined with U-Net encoder and contrast learning technology, an image quality evaluation model is built to improve the feature extraction ability and interpretability of the model.

Benefits of technology

It improves the accuracy and reliability of image quality evaluation, enhances the interpretability of the model, and allows users to better understand the basis for evaluation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119941691A_ABST
    Figure CN119941691A_ABST
Patent Text Reader

Abstract

The invention discloses an image quality evaluation method based on data enhancement and contrast learning, and belongs to the technical field of image quality evaluation and deep learning. The method comprises the following steps: preprocessing an original image, and inputting the preprocessed image into an image quality evaluation model to obtain an image quality score; the preprocessing module is composed of random data enhancement, depth image prior and bilinear interpolation; the image quality evaluation model is constructed by positive and negative sample pairs, a feature extraction module, and a network training and score generation module. The feature extraction module adopts a U-Net encoder and comprises a down-sampling layer and a convolution layer; the network training comprises NT-Xent loss function calculation and paired sample training; the score generation module is used for predicting and obtaining an image quality score by using a regression model; according to the method, the network is constructed based on strict mathematical modeling, so that the model has rigorous interpretability, and the performance and reliability of image quality evaluation can be improved based on theoretical derivation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to an image quality assessment method based on data enhancement and contrast learning, and belongs to the technical field of image quality evaluation and deep learning. Background Art

[0002] With the rapid development of the Internet and digital media, the generation, processing and dissemination of images have become more and more common. The assessment of image quality is not only crucial to industries such as visual arts, advertising, and social media, but also plays an important role in medical imaging, surveillance systems, and autonomous driving. High-quality images can provide clearer information and help users make more informed decisions.

[0003] Traditional image quality assessment methods mainly fall into two categories: subjective assessment and objective assessment. Subjective assessment relies on feedback from human observers, usually using rating questionnaires and laboratory tests to collect user opinions. However, this method is not only time-consuming, but the results are often affected by individual differences between observers [1]. Objective assessment methods automatically generate quality scores through computer algorithms, such as peak signal-to-noise ratio, structural similarity index, etc. [2]. Although these methods can quickly provide scores, they often cannot fully reflect human subjective perception and have certain limitations.

[0004] In recent years, the rise of deep learning technology has provided new opportunities for image quality assessment. Convolutional neural networks (CNNs) have been widely used in various visual tasks due to their powerful ability in image feature extraction. By training CNN models, researchers can capture high-level features in images, thereby achieving more accurate quality assessment[3]. In addition, contrastive learning, as an unsupervised learning method, can effectively improve the quality of feature representation by modeling the similarities between samples, providing a new idea for image quality assessment[4].

[0005] Although some deep learning-based image quality assessment methods have been proposed, these methods still have shortcomings in handling various types of distortion and in different scenarios. In addition, existing models often lack interpretability, making it difficult for users to understand the evaluation basis of the model. Therefore, it is particularly important to explore more comprehensive image quality assessment methods that can both improve the accuracy of the assessment and enhance the interpretability of the model.

[0006] In summary, the image quality assessment method based on convolutional neural network, data enhancement technology and contrastive learning aims to fill the gap in the prior art and improve the performance and reliability of image quality assessment by integrating multiple features and learning strategies.

[0007] Smith, John, and Rachel Lee. "Evaluating Multimedia Quality Using the Mean Opinion Score." *Journal of Multimedia Research*, vol. 12, no. 3, 2020,pp. 45-60. (Smith, John, Rachel Lee. "Evaluating Multimedia Quality Using the Mean Opinion Score." Journal of Multimedia Research, Vol. 12, No. 3, 2020, pp. 45-60.) Wang, Z., and A.C. Bovik. "A Universal Image Quality Index." *IEEE Signal Processing Letters*, vol. 11, no. 8, 2004, pp. 637-640. Liu, J., Zhou, L. (2021). Application of self-supervised contrastive learning in blind image quality assessment. Journal of Electronics, 49(8), 1644-1652.

[0008] Zhou, B.,&Shi, Z. (2023). Deep Contrastive Learning for Image Quality Evaluation and Enhancement. Signal Processing: Image Communication, 98, 116-128.(Zhou Bo, Shi Zhi (2023). Deep Contrastive Learning for Image Quality Evaluation and Enhancement. Signal Processing: Image Communication, 98, 116-128.) Summary of the invention

[0009] The purpose of the present invention is to overcome the shortcomings of the prior art and provide an image quality assessment method based on data enhancement and contrastive learning to solve the problems of large sample requirements, large model parameter scale and high training difficulty of the current deep learning image quality assessment algorithm; In order to achieve the above purpose / solve the above technical problems, the present invention is implemented by adopting the following technical solutions: An image quality assessment method based on data enhancement and contrastive learning, comprising: Preprocess the image, and input the preprocessed image into a pre-built image quality evaluation model to obtain the image quality score; Among them, the preprocessing includes random data enhancement, deep image prior, and bilinear interpolation operations on the image; the image quality assessment model includes a feature extraction module, a network training and a score generation module, the feature extraction module is used to extract high-level features of the constructed positive and negative sample pairs, the network training is trained through images and high-level features, and the score generation module is used to predict the quality score of the image based on the high-level feature vector.

[0010] Optionally, the deep image prior is an untrained convolutional neural network that is initialized first. The network architecture consists of a convolutional layer, an activation function, and a deconvolutional layer. The initial parameters of the network architecture are randomly initialized. The convolutional neural network is used as an optimization tool to directly repair the image. The input of the network architecture is a random noise image z, and the output is the repaired image. ; By minimizing the loss function, the weights and biases of the network are optimized so that the output image gradually becomes a real image during the optimization process. First, the data loss is calculated. The operation is: (1) in, represents the pixel positions of all known regions, represents the restored image, represents the original image; Then calculate the regularization loss to constrain the structure of the generated image. The operation is: (2) in, is an image At pixel position The gradient indicates the degree of pixel change. It means to sum or calculate each pixel x in the image, x represents each pixel position in the image, and all refers to all pixels in the image; The final loss function is the weighted sum of data loss and regularization loss, defined as: +λ (3) Among them, λ is a hyperparameter that controls the balance between data loss and regularization loss; The gradient of the loss function for each parameter is calculated, the calculated gradient is back-propagated, and the parameters of the network are updated. The back-propagation calculates the gradient of each parameter by the chain rule, and is passed back from the last layer to the input layer until the loss function converges or a predetermined maximum number of training times is reached.

[0011] Optionally, the bilinear interpolation is used to perform secondary restoration on the image, first identifying the missing areas in the depth image, and then To express, Indicates that the position is missing. Indicates that the position is known, for each missing point , select its left, right, top and bottom neighboring points for bilinear interpolation, where horizontal interpolation is expressed as: (4) (5) is fixed, and Respectively expressed in and The linear interpolation result on ; the vertical interpolation can be expressed as: (6) The depth value of the missing area Updated to the interpolated value.

[0012] Optionally, the construction of the positive and negative sample pairs includes: The preprocessed images are matched one-to-one with the images generated by the random data augmentation and bilinear interpolation method to construct positive sample pairs; The images obtained by rotation, scaling and bilinear interpolation are first randomly shuffled, and the image pairs after shuffle transformation are marked as negative sample pairs; Among them, after marking the positive and negative sample pairs, the image is normalized. The normalization uses scale normalization to scale all sample data to the same scale and distribute them symmetrically about 0. The specific normalization function is: (7) in, is the normalized image, is the input noisy image.

[0013] Optionally, the feature extraction module adopts a U-Net encoder, and uses the U-Net encoder to convolve and downsample the input image to achieve the purpose of feature extraction; a difficult negative sample mining method is additionally used for negative samples, and difficult negative samples in the training process are selected to promote the model to learn more discriminative feature representations.

[0014] Optionally, the U-Net encoder includes a convolution layer and downsampling. The convolution operation extracts local features of the image. A 3×3 convolution kernel is used. Assuming that the input image is , then the convolution operation can be expressed as: (8) in is the output feature map at position The value of is the pixel value of the input image, is the convolution kernel at position The value of Downsampling is done through maximum pooling and average pooling, where maximum pooling reduces the dimension of the feature map by selecting the maximum value in the pooling window. Using a 2×2 pooling window, the maximum pooling operation is expressed by the following formula: (9) Average pooling performs dimensionality reduction by selecting the average of all values ​​in the pooling window. Similar to maximum pooling, using a 2×2 pooling window, the average pooling operation is expressed as: (10) After completing the maximum pooling and average pooling, the outputs are concatenated using channel-wise merging, and the concatenation is expressed as: (11) Convolution and pooling should be performed alternately to extract more and more abstract features layer by layer.

[0015] Optionally, the selecting of difficult negative samples in the training process includes: for the positive and negative sample pairs , first calculate the cosine similarity of the positive and negative sample pairs, the corresponding formula is: (12) in, and The samples are and The characteristic vector of represents the dot product, Represents the Euclidean norm, for positive samples , select the negative sample with the smallest distance from the negative sample set ,in is a set of negative sample pairs, and calculates their similarity: (13) Negative samples with high similarity are difficult negative samples; NT-Xent loss is used to focus on distinguishing difficult negative samples. The loss function formula is: (14) in, is a positive sample, is a negative sample, is a hard negative sample, is the temperature coefficient.

[0016] Optionally, the network training is performed by adjusting the training process through the following steps, including: Step S1, select a pair of image samples from the training data and calculate their distance metric in the embedding space by cosine similarity; Step S2, based on the similarity or difference of paired samples, the performance of the model is evaluated by calculating the NT-Xent loss function. For each positive sample pair, their similarity is maximized, while for the negative sample pair, the similarity of all samples in the denominator is compared to ensure that the similarity of the negative sample is minimized; Step S3, by adjusting the temperature parameters Scale the similarity, which in turn affects the calculation process of cosine similarity. Temperature parameter The smoothness of the loss function is controlled. A lower temperature will make the model focus more on the differences between similar sample pairs, while a higher temperature will increase the influence of negative sample pairs, resulting in a smoother learning process for the model.

[0017] Optionally, the score generation module predicts the quality score of the image according to the high-level feature vector extracted from the U-Net, using a fully connected regression network, including an input layer, multiple fully connected layers and an output layer, wherein the input of the input layer is the high-level feature vector extracted by the U-Net, and its shape should be the same as the dimension of the feature vector; Assume the fully connected input is , the output of this layer It can be calculated by the following formula: (15) in It is The weight matrix of the layer, It is The bias vector of the layer, yes Activation function, assuming that the output of the fully connected layer is , the calculation formula of the output layer is: (16) is the quality score of the image, is the weight matrix of the output layer, is the bias vector of the output layer.

[0018] Optionally, the training of the score generation module includes: Step SS1, importing a dataset containing images and their corresponding true quality scores; Step SS2, normalize all images and input them into the feature extraction module for feature extraction; Step SS3, predicting the quality score of the image; In step SS4, the mean square error (MSE) is used as the loss function to measure the difference between the model's predicted score and the actual score. The corresponding formula is: (17) in, is the true subjective rating, is the rating predicted by the model, and N is the size of the dataset; Step SS5: Back propagate the error of each layer and calculate the gradient of each layer parameter ; Step SS6, choose an appropriate learning rate , according to the calculated gradient, update the parameters according to the gradient descent rule, iterate multiple times until the loss function converges, and the update formula is: .

[0019] Compared with the prior art, the present invention has the following beneficial effects: The present invention provides a method for generating training samples through data enhancement, extracting features using CNN, and evaluating image quality using contrastive learning. The model's ability to judge image quality is enhanced through multi-level feature extraction and high-level feature calculation. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] Figure 1 is a flow chart of a method according to an embodiment of the present invention; Figure 2 This is a structural diagram of an image preprocessing module according to an embodiment of the present invention; Figure 3 This is a structural diagram of a positive and negative sample pair construction and feature extraction module according to an embodiment of the present invention; Figure 4 A diagram showing a network training structure according to an embodiment of the present invention; Figure 5 This is a structural diagram of a scoring generation module according to an embodiment of the present invention; Figure 6 This is a diagram of the fully connected regression network training structure of an embodiment of the present invention. DETAILED DESCRIPTION

[0021] In order to make the technical means, creative features, objectives and effects achieved by the present invention easy to understand, the present invention is further explained below in conjunction with specific implementation methods.

[0022] In the description of the present invention, it should be understood that the terms "center", "longitudinal", "lateral", "up", "down", "front", "back", "left", "right", "vertical", "horizontal", "top", "bottom", "inside", "outside" and the like indicate positions or positional relationships based on the positions or positional relationships shown in the accompanying drawings, and are only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore cannot be understood as limiting the present invention. In addition, the terms "first", "second", and the like are only used for descriptive purposes, and cannot be understood as indicating or implying relative importance or implicitly indicating the number of technical features indicated. Thus, features defined as "first", "second", and the like may explicitly or implicitly include one or more of the features. In the description of the present invention, unless otherwise specified, "multiple" means two or more.

[0023] In the description of the present invention, it should be noted that, unless otherwise clearly specified and limited, the terms "installed", "connected", and "connected" should be understood in a broad sense, for example, it can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be a direct connection, or it can be indirectly connected through an intermediate medium, or it can be the internal communication of two components. For ordinary technicians in this field, the specific meanings of the above terms in the present invention can be understood by specific circumstances.

[0024] like Figure 1-Figure 6 As shown, the present invention discloses an image quality assessment method based on data enhancement and contrast learning, comprising: Step 1: Preprocess the reference image and the distorted image; Step 2, input the preprocessed image into a pre-built image quality assessment model to obtain an image quality score; Please see Figure 2 ,In one embodiment, the preprocessing includes random data enhancement, deep image prior, and bilinear interpolation; In one embodiment, the random data enhancement includes: first performing a series of random data enhancement operations on the input image to generate more training samples. These enhancement operations include but are not limited to: random flipping, random cropping, random rotation, random scaling and stretching, and color conversion. These enhancement operations are independently applied to each image in a probabilistic manner to generate more diverse training samples.

[0025] In one embodiment, in step 1, preprocessing the image includes: the deep image is initialized with an untrained convolutional neural network. The network architecture consists of a convolutional layer, an activation function, and a deconvolutional layer. The initial parameters of the network are randomly initialized. The convolutional neural network (CNN) is used as a tool in the optimization process to directly repair the image without any previous image data set for pre-training. The input of the network is a random noise image z, and the output is the repaired image ; By minimizing the loss function, the weights and biases of the network are optimized so that the output image gradually becomes close to the real image during the optimization process. First, the data loss is calculated. The operation is: (1) in, represents the pixel positions of all known regions, Represents the original image.

[0026] Then calculate the regularization loss to constrain the structure of the generated image and prevent overfitting or generating unreasonable images. The operation is: (2) in, is an image At pixel position The gradient indicates the degree of pixel change. For each pixel in the image To perform a sum or calculation. Represents each pixel position in the image, and "all" refers to all pixels in the image.

[0027] The final loss function is the weighted sum of data loss and regularization loss, defined as: +λ (3) Where λ is a hyperparameter that controls the balance between data loss and regularization loss.

[0028] Finally, the gradient of the loss function for each parameter (weight and bias) is calculated, and these gradients are back-propagated to update the network parameters. Back-propagation is to calculate the gradient of each parameter through the chain rule, and pass it backwards from the last layer to the input layer. The Adam optimizer dynamically adjusts the learning rate according to the gradient history of each parameter and updates the parameters; the above process is repeated until the loss function converges (that is, reaches the minimum value) or reaches the predetermined maximum number of training times.

[0029] In one embodiment, in step 1, the bilinear interpolation is used to perform secondary restoration on the image. First, the missing area in the depth image is identified, and the mask is used to To express, Indicates that the position is missing. Indicates that the position is known. For each missing point , select its left, right, top and bottom neighboring points for bilinear interpolation. The horizontal interpolation is expressed as: (4) (5) is fixed, and Respectively expressed in and The linear interpolation result on .

[0030] Vertical interpolation can be expressed as: (6) The depth value of the missing area Update to the interpolated value; Please see Figure 3 In one embodiment, in step 2, the positive and negative sample pairs are constructed using the output images of the three steps of preprocessing, and the positive sample pairs are constructed by matching the random data enhancement with the images generated by the bilinear interpolation method. The difference between the two is small, so there is a high similarity in the image content, which meets the characteristics of the positive sample. The images obtained by rotation, scaling and bilinear interpolation are first randomly shuffled. Since these images have large structural differences and pixel-level changes, the generated images have large visual and structural differences from the original images, so these transformed image pairs are marked as negative sample pairs.

[0031] In one embodiment, the positive and negative sample pairs are constructed, and the image is normalized after the positive and negative sample pairs are marked. The normalization uses scale normalization to scale all sample data to the same scale and is symmetrically distributed about 0. The specific normalization function is: (7) in, is the normalized image, is the input noisy image.

[0032] In one embodiment, in step 2, the feature extraction module adopts a U-Net encoder, and uses the U-Net encoder to convolve and downsample the input image to achieve the purpose of feature extraction; the hard negative mining technology is additionally used for negative samples, and the "difficult" negative samples in the training process are selected to promote the model to learn more discriminative feature representations.

[0033] In one embodiment, in step 2, the encoder includes a convolution layer and downsampling, and the convolution operation extracts local features of the image using a 3×3 convolution kernel. Assume that the input image is , then the convolution operation can be expressed as: (8) in is the output feature map at position The value of . is the pixel value of the input image, is the convolution kernel at position The value of .

[0034] Downsampling is done through maximum pooling and average pooling. The maximum pooling reduces the dimension of the feature map by selecting the maximum value in the pooling window. Using a 2×2 pooling window, the maximum pooling operation can be expressed by the following formula: (9) Average pooling performs dimensionality reduction by selecting the average of all values ​​in the pooling window. Similar to max pooling, using a 2×2 pooling window, the average pooling operation can be expressed as: (10) After completing the maximum pooling and average pooling, their outputs are concatenated together using channel-wise merging so that the features of both pooling operations are retained and fed into the next layer. Their concatenation can be expressed as: (11) Convolution and pooling should be performed alternately to extract more and more abstract features layer by layer; In one embodiment, in step 2, the difficult negative sample mining is performed for the positive and negative sample pairs. , first calculate the cosine similarity of the positive and negative sample pairs, the corresponding formula is (12) in, and The samples are and The characteristic vector of represents the dot product, represents the Euclidean norm. For positive samples , select the negative sample with the smallest distance from the negative sample set ,in is a set of negative sample pairs), and calculate their similarity: (13) This negative sample with a high similarity is called a difficult negative sample.

[0035] In one embodiment, in step 2, the loss function calculation can be performed by adjusting the loss function so that it pays more attention to the distinction of these difficult negative samples, using NT-Xent loss; the loss function formula is: (14) in, is a positive sample, is a negative sample, is a hard negative sample, is the temperature coefficient.

[0036] Please see Figure 4 In one embodiment, in step 2, the network training is adjusted by the following steps, and the training process includes the following steps: Step S1, select a pair of image samples from the training data, which can be a positive sample pair or a negative sample pair, and calculate their distance metric in the embedding space by cosine similarity; Step S2, evaluate the performance of the model by calculating the NT-Xent loss function based on the similarity or difference of paired samples. For each positive sample pair, their similarity is maximized. For the negative sample pair, the similarity of all samples in the denominator is used for comparison to ensure that the similarity of the negative samples is minimized.

[0037] Step S3, temperature parameters Scale the similarity, which in turn affects the calculation process of cosine similarity. Temperature parameter Controls the smoothness of the loss function. A lower temperature will make the model focus more on the differences between similar pairs of samples, while a higher temperature will increase the influence of negative pairs, resulting in a smoother learning process for the model.

[0038] Please see Figure 5 In one embodiment, in step 2, the score generation module predicts the quality score of the image based on the high-level feature vector extracted from U-Net. For this purpose, a fully connected regression network is used, including an input layer, multiple fully connected layers and an output layer. The input of the input layer is the high-level feature vector extracted by U-Net, and its shape should be the same as the dimension of the feature vector.

[0039] Assume the fully connected input is , the output of this layer It can be calculated by the following formula: (15) in It is The weight matrix of the layer, It is The bias vector of the layer, yes Activation function. Assume that the output of the fully connected layer is (where L is the number of the last layer), the calculation formula of the output layer is: (16) is the quality score of the image, is the weight matrix of the output layer, is the bias vector of the output layer; Please see Figure 6 In one embodiment, in step 2, the score generation module is trained by the following steps: Step SS1, importing a dataset containing images and their corresponding true quality scores; The image and image quality score dataset uses TID2013; Step SS2, normalize all images and input them into the feature extraction module for feature extraction; The feature extraction method is the same as the feature extraction method described in step 2 to ensure the consistency of the training environment and the application environment; Step SS3, predicting the quality score of the image; The rating prediction method is the same as the rating generation module described in step 2; In step SS4, the mean square error (MSE) is used as the loss function to measure the difference between the model predicted score and the actual score. The corresponding formula is: (17) in, is a true subjective rating. is the score predicted by the model, and N is the size of the dataset; Step SS5: Back propagate the error of each layer and calculate the gradient of each layer parameter ; Step SS6, choose an appropriate learning rate , according to the calculated gradient, update the parameters according to the gradient descent rule. Iterate multiple times until the loss function converges. The update formula is:

[0040] The specific implementation method of each unit is the same as each step and will not be described in detail in the present invention.

[0041] The above is only a preferred embodiment of the present invention. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the technical principles of the present invention. These improvements and modifications should also be regarded as the scope of protection of the present invention.

Claims

1. An image quality assessment method based on data enhancement and contrastive learning, characterized in that: include: Preprocess the image, and input the preprocessed image into a pre-built image quality evaluation model to obtain the image quality score; Among them, the preprocessing includes random data enhancement, deep image prior, and bilinear interpolation operations on the image; the image quality assessment model includes a feature extraction module, a network training and a score generation module, the feature extraction module is used to extract high-level features of the constructed positive and negative sample pairs, the network training is trained through images and high-level features, and the score generation module is used to predict the quality score of the image based on the high-level feature vector.

2. The image quality assessment method based on data enhancement and contrastive learning according to claim 1, characterized in that: The deep image prior is to initialize an untrained convolutional neural network. The network architecture consists of a convolutional layer, an activation function and a deconvolutional layer. The initial parameters of the network architecture are randomly initialized. The convolutional neural network is used as an optimization tool to directly repair the image. The input of the network architecture is a random noise image z, and the output is the repaired image. ; By minimizing the loss function, the weights and biases of the network are optimized so that the output image gradually becomes a real image during the optimization process. First, the data loss is calculated. The operation is: (1) in, represents the pixel positions of all known regions, represents the restored image, represents the original image; Then calculate the regularization loss to constrain the structure of the generated image. The operation is: (2) in, is an image At pixel position The gradient indicates the degree of pixel change. It means to sum or calculate each pixel x in the image, x represents each pixel position in the image, and all refers to all pixels in the image; The final loss function is the weighted sum of data loss and regularization loss, defined as: +λ (3) Among them, λ is a hyperparameter that controls the balance between data loss and regularization loss; The gradient of the loss function for each parameter is calculated, the calculated gradient is back-propagated, and the parameters of the network are updated. The back-propagation calculates the gradient of each parameter by the chain rule, and is passed back from the last layer to the input layer until the loss function converges or a predetermined maximum number of training times is reached.

3. The image quality assessment method based on data enhancement and contrastive learning according to claim 1, characterized in that: The bilinear interpolation is used to perform secondary restoration on the image. First, the missing area in the depth image is identified, and then the missing area is restored by masking. To express, Indicates that the position is missing. Indicates that the position is known, for each missing point , select its left, right, top and bottom neighboring points for bilinear interpolation, where horizontal interpolation is expressed as: (4) (5) is fixed, and Respectively expressed in and The linear interpolation result on ; the vertical interpolation can be expressed as: (6) The depth value of the missing area Updated to the interpolated value.

4. The image quality assessment method based on data enhancement and contrastive learning according to claim 1, characterized in that: The construction of the positive and negative sample pairs includes: The preprocessed images are matched one-to-one with the images generated by the random data augmentation and bilinear interpolation method to construct positive sample pairs; The images obtained by rotation, scaling and bilinear interpolation are first randomly shuffled, and the image pairs after shuffle transformation are marked as negative sample pairs; Among them, after marking the positive and negative sample pairs, the image is normalized. The normalization uses scale normalization to scale all sample data to the same scale and distribute them symmetrically about 0. The specific normalization function is: (7) in, is the normalized image, is the input noisy image.

5. The image quality assessment method based on data enhancement and contrastive learning according to claim 1, characterized in that: The feature extraction module adopts a U-Net encoder, and uses the U-Net encoder to convolve and downsample the input image to achieve the purpose of feature extraction; a difficult negative sample mining method is additionally used for negative samples, and difficult negative samples in the training process are selected to promote the model to learn more discriminative feature representations.

6. The image quality assessment method based on data enhancement and contrastive learning according to claim 5, characterized in that: The U-Net encoder includes convolutional layers and downsampling. The convolution operation extracts local features of the image. A 3×3 convolution kernel is used. Assuming the input image is , then the convolution operation can be expressed as: (8) in is the output feature map at position The value of is the pixel value of the input image, is the convolution kernel at position The value of Downsampling is done through maximum pooling and average pooling, where maximum pooling reduces the dimension of the feature map by selecting the maximum value in the pooling window. Using a 2×2 pooling window, the maximum pooling operation is expressed by the following formula: (9) Average pooling performs dimensionality reduction by selecting the average of all values ​​in the pooling window. Similar to maximum pooling, using a 2×2 pooling window, the average pooling operation is expressed as: (10) After completing the maximum pooling and average pooling, the outputs are concatenated using channel-wise merging, and the concatenation is expressed as: (11) Convolution and pooling should be performed alternately to extract more and more abstract features layer by layer.

7. The image quality assessment method based on data enhancement and contrastive learning according to claim 5, characterized in that: The difficult negative samples in the selection training process include: for the positive and negative sample pairs , first calculate the cosine similarity of the positive and negative sample pairs, the corresponding formula is: (12) in, and The samples are and The characteristic vector of represents the dot product, Represents the Euclidean norm, for positive samples , select the negative sample with the smallest distance from the negative sample set ,in is a set of negative sample pairs, and calculates their similarity: (13) Negative samples with high similarity are difficult negative samples; NT-Xent loss is used to focus on distinguishing difficult negative samples. The loss function formula is: (14) in, is a positive sample, is a negative sample, is a hard negative sample, is the temperature coefficient.

8. The image quality assessment method based on data enhancement and contrastive learning according to claim 1, characterized in that: The network training is performed by adjusting the training process through the following steps, including: Step S1, select a pair of image samples from the training data and calculate their distance metric in the embedding space by cosine similarity; Step S2, based on the similarity or difference of paired samples, the performance of the model is evaluated by calculating the NT-Xent loss function. For each positive sample pair, their similarity is maximized, while for the negative sample pair, the similarity of all samples in the denominator is compared to ensure that the similarity of the negative sample is minimized; Step S3, by adjusting the temperature parameters Scale the similarity, which in turn affects the calculation process of cosine similarity. Temperature parameter The smoothness of the loss function is controlled. A lower temperature will make the model focus more on the differences between similar sample pairs, while a higher temperature will increase the influence of negative sample pairs, resulting in a smoother learning process for the model.

9. The image quality assessment method based on data enhancement and contrastive learning according to claim 1, characterized in that: The score generation module predicts the quality score of the image based on the high-level feature vector extracted from the U-Net, using a fully connected regression network, including an input layer, multiple fully connected layers, and an output layer. The input of the input layer is the high-level feature vector extracted by the U-Net, and its shape should be the same as the dimension of the feature vector; Assume the fully connected input is , the output of this layer It can be calculated by the following formula: (15) in It is The weight matrix of the layer, It is The bias vector of the layer, yes Activation function, assuming that the output of the fully connected layer is , the calculation formula of the output layer is: (16) is the quality score of the image, is the weight matrix of the output layer, is the bias vector of the output layer.

10. The image quality assessment method based on data enhancement and contrastive learning according to claim 1, characterized in that: The training of the score generation module includes: Step SS1, importing a dataset containing images and their corresponding true quality scores; Step SS2, normalize all images and input them into the feature extraction module for feature extraction; Step SS3, predicting the quality score of the image; In step SS4, the mean square error (MSE) is used as the loss function to measure the difference between the model's predicted score and the actual score. The corresponding formula is: (17) in, is the true subjective rating, is the rating predicted by the model, and N is the size of the dataset; Step SS5: Back propagate the error of each layer and calculate the gradient of each layer parameter ; Step SS6, choose an appropriate learning rate , according to the calculated gradient, update the parameters according to the gradient descent rule, iterate multiple times until the loss function converges, and the update formula is: .

Citation Information

Patent Citations

  • Enhanced image quality evaluation method based on twin network

    CN110033446A

  • Method for repairing texture of terracotta figure by combining recursive reasoning and feature re-optimization

    CN116309157A

  • Non-reference image quality evaluation method based on multi-domain distortion learning

    CN116823794A

  • Self-supervised reference-free image quality evaluation method based on semantic comprehension

    CN117351002A

  • Weak supervision image semantic segmentation method based on sample mixing and contrast learning

    CN118154884A

Cited By

  • Radiotherapy knowledge intelligent question answering and guide updating method and system with feedback optimization mechanism

    CN121434221A

  • A method and system for intelligent Q&A and guideline updates for radiotherapy knowledge with feedback optimization mechanisms

    CN121434221B