Low-bit-rate semi-reference image quality inspection method and system based on self-distillation
By using self-distillation technology to generate low-bitrate reference streams in smart IoT scenarios, and combining the dual-head structure of the teacher network and the student network, the problems of high-precision and real-time image quality assessment under extremely low bitrate conditions are solved, and efficient image quality assessment on edge devices is achieved.
Patent Information
- Application Number
- CN202511086423.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-05
- Publication Date
- 2025-09-09
- Estimated Expiration
- 2045-08-05
AI Technical Summary
In the smart IoT scenario, existing image quality assessment technologies find it difficult to obtain high-precision reference information under extremely low bitrate conditions, and are unable to complete quality assessment in real time on chips with limited computing power. In particular, the prediction accuracy is insufficient in multi-distortion mixed scenarios.
Self-distillation technology is used to extract image semantic vectors by freezing the basic image encoder network model, and product quantization is performed to generate a low-rate reference bitstream. The twin-head structure combining the teacher network and the student network is used to generate pseudo reference images and align cross-image differences to achieve high-precision quality assessment.
It reduces the bandwidth occupancy of reference data, improves the evaluation accuracy of multi-distortion mixed scenarios, and enables real-time computing on edge devices to meet the real-time requirements of smart IoT terminals.
Smart Images

Figure CN120612552A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of image quality inspection technology, and in particular to a low-bitrate semi-reference image quality inspection method and system based on self-distillation. Background Art
[0002] In smart IoT scenarios, intelligent visual terminals (such as NB-IoT security cameras, remote agricultural micro-cameras, and tower inspection cameras) are typically deployed in edge environments with limited uplink bandwidth and computing resources. In these scenarios, terminals need to continuously report images or video streams, and the backend must monitor image quality in real time to dynamically adjust encoding bitrates, trigger error retransmissions, or issue alarms.
[0003] Existing image quality assessment technologies have obvious shortcomings in the smart IoT environment: full-reference image quality assessment relies on the complete original image for comparison, and requires the additional transmission of hundreds of KB or even MB of lossless or high-quality reference images, which exceeds the carrying capacity of the IoT link; non-reference image quality assessment is based only on the characteristics of the distorted image itself. Although it occupies the least bandwidth, it is difficult to guarantee the prediction accuracy when facing multiple distortion mixed scenarios such as compression artifacts, random noise, motion blur and color drift; the existing semi-reference assessment method uses low-level statistical features such as histograms and edges. Although the reference bit rate is low, it is not adaptable enough to structure-color mixing distortion; at the same time, the number of network parameters for difference modeling generally exceeds 20M, which cannot be run in real time on low-power MCUs or edge gateways.
[0004] Therefore, how to obtain reference information sufficient to support high-precision evaluation under extremely low bit rate conditions and complete quality evaluation in real time on chips with limited computing power has become a core technical problem that needs to be urgently solved in smart IoT smart terminals. Summary of the Invention
[0005] This application provides a low-bitrate semi-reference image quality inspection method and system based on self-distillation to solve the problem of how to obtain reference information sufficient to support high-precision evaluation under extremely low bitrate conditions and complete quality evaluation in real time on a chip with limited computing power.
[0006] In a first aspect, the present application provides a low-bitrate semi-reference image quality inspection method based on self-distillation, the method comprising: Acquire an image sample set of an intelligent visual terminal and perform a normalization process to obtain a normalized image; wherein the image sample set is composed of a plurality of image samples; A frozen basic image encoder network model is used to extract the image semantic vector of the normalized image; Perform product quantization on the image semantic vector to obtain a reference bitstream; The distorted image and reference bitstream corresponding to the image sample are synchronously sent to the receiving end, and the receiving end obtains the reconstructed semantic vector corresponding to the image sample according to the reference bitstream; Generate a pseudo reference image based on the reconstructed semantic vector and the distorted image; Perform cross-image difference alignment on the pseudo reference image and the distorted image in a preset salient region to obtain a difference feature tensor; Based on the difference feature tensor, the teacher network obtains inputs to the quality regression sub-head and the distortion classification sub-head. Using the difference feature tensor, the teacher network obtains a quality prediction confidence index and a distortion classification prediction confidence index for each image sample in the image sample set, thereby determining the type of each image sample. The types are categorized as low uncertainty, medium uncertainty, and high uncertainty. Based on the difference feature tensor and the type of each sample, the teacher network parameters are frozen to obtain a trained student network. Obtain the difference feature tensor of the test data, use the trained student network to output the subjective quality score and preset several types of distortion probability vectors ,Pick is the dominant distortion type of the test data.
[0007] In one implementation of the present application, obtaining a sample set of an intelligent visual terminal and performing a normalization process to obtain a normalized image specifically includes: Terminal camera collects intelligent visual terminal sample set , the resolution is fixed at 1280*720 pixels and the frame format is YUV-420; right Apply 3×3 Gaussian filter kernel to the three chroma components respectively; Apply histogram equalization to adjust the brightness distribution to make the pixel grayscale uniform; Convert the processed image to RGB and normalize it to the range [-1, 1] according to the following formula.
[0008] In one implementation of the present application, a frozen basic image encoder network model is used to extract the image semantic vector of the normalized image, specifically including: The normalized image Feed into the basic image encoder; through the formula: , read the last layer CLS Token of the basic image encoder network to obtain the image semantic vector ; in, is the frozen basic image encoder network model; Represents the i-th semantic dimension.
[0009] In one implementation of the present application, product quantization is performed on the image semantic vector to obtain a reference bitstream, specifically including: According to the preset division formula: , divide the image semantic vector into m=16 32-dimensional sub-vectors in turn ; According to the preset formula: , for each In the corresponding codebook Calculate the Euclidean distance and obtain the first-level quantitative index ; Where K represents the number of codewords in each codebook; =[ ] , represents the mth codebook; {1,...,256} represents the first-level quantization index; Save the 16 first-level quantization indexes in sequence as a 16-Byte unsigned integer array to obtain the first-level quantization reconstruction vector ; Perform the operation on each sub-residual: ;in, represents the secondary residual index, represents the first-level quantized residual of the m-th sub-vector, represents the jm-th codeword in the m-th residual codebook; The first-level quantization index of each sub-vector and is the secondary residual index The index pairs are combined into index pairs, and the index pairs are compressed into reference bitstream b through RangeCoder entropy coding.
[0010] In one implementation of the present application, the receiving end obtains the reconstructed semantic vector according to the reference bitstream, specifically including: The receiving end obtains reference stream b: By formula: , decode the reference bitstream into a reconstructed semantic vector ; Represents the decoding function.
[0011] In one implementation of the present application, generating a pseudo reference image based on the reconstructed semantic vector and the distorted image specifically includes: The semantic vector will be reconstructed Input mapping network , output text prompt t.
[0012] Distorted image As the starting noise, load the Latent Diffusion model; Perform 5 steps of DPM-Solver sampling to obtain a pseudo reference image ;in, Represents the 5th step DPM-Solver reverse sampler of the LatentDiffusion model, ∈ represents the pseudo reference image.
[0013] In one implementation of the present application, cross-image difference alignment is performed on the pseudo reference image and the distorted image in a preset salient region to obtain a difference feature tensor, specifically including: The pseudo reference image and the distorted image are divided into 16×16 patches, and the token sequence is obtained as follows: ; By formula: , calculate the self-attention weight ; And select the top-20% patches to form the set β; in, represents the average attention of the i-th patch; , represents the h-th head query / key weight matrix; Input 6 layers of Cross-ViT to the token in β and output the difference feature tensor D.
[0014] In one implementation of the present application, the teacher network uses the difference feature tensor to obtain the quality prediction confidence index and distortion classification prediction confidence index of each sample in the sample set, and then determines the type of each sample, specifically including: For the same difference feature tensor D, Dropout is enabled in the teacher network and M = 3 forward propagations are performed to simulate the prediction volatility of the network output, obtaining 3 sets of quality prediction results {Q(1), Q(2), Q(3)} and the corresponding distortion type classification probability vectors {c(1), c(2), c(3)}; Calculate the variance of the quality prediction scores three times and get , as a quality prediction confidence indicator; By formula: , calculate the entropy value of the three classification probability vectors It will serve as a confidence indicator for the distortion classification prediction; According to the quality prediction confidence index and the distortion classification prediction confidence index of all samples, the quality prediction confidence index distribution and the distortion classification prediction confidence index distribution are obtained; According to the preset quantiles of the distribution, the quality prediction volatility threshold and the classification prediction confidence threshold are obtained; When the quality prediction confidence index of a sample is less than or equal to the quality prediction volatility threshold, and the distortion classification prediction confidence index is less than or equal to the classification prediction confidence threshold, the sample is of low uncertainty type; When the quality prediction confidence index of a sample is greater than the quality prediction volatility threshold or the distortion classification prediction confidence index is greater than the classification prediction confidence threshold, the sample is of medium uncertainty type; When the quality prediction confidence index of a sample is greater than the quality prediction volatility threshold and the distortion classification prediction confidence index is greater than the classification prediction confidence threshold, the sample is of high uncertainty type.
[0015] In one implementation of the present application, based on the difference feature tensor and the type of each sample, the student network is used to output the subjective quality score and several preset distortion probability vectors. ,Pick The dominant distortion types include: Input the difference feature tensor, after mean pooling and maximum pooling, we get: ;in, () is the mean pooling function, () is the maximum pooling function; based on , perform two-layer full connection on the quality sub-head, dimension change: 384 → 64 → 1, output subjective quality prediction ; based on , the classification sub-head performs two-layer full connection, the dimension changes to: 384 → 64 → 8, and the distortion probability is obtained by Softmax ;in, ; According to the formula: , calculate the uncertainty weighted distillation loss; in, 、 The subjective quality prediction and distortion probability of the teacher network output; in , The weights are set as follows: , , in, Indicates the preset adjustment factor; Backpropagation only updates the student network and the quantized codebook; the teacher network is always frozen; Access to a network of trained students; Input the difference feature tensor into the trained student network and output the subjective quality score and distortion probability vector ,Pick is the dominant distortion type.
[0016] In a second aspect, the present application provides a low-bitrate semi-reference image quality inspection system based on self-distillation, the system comprising: An image acquisition module is used to acquire a sample set of an intelligent visual terminal and perform standardization processing to obtain a normalized image; A vector extraction module is used to extract normalized image semantic vectors using a frozen basic image encoder network model; A code stream acquisition module is used to perform product quantization on the image semantic vector to obtain a reference code stream; The vector reconstruction module is used to synchronously send the distorted image and reference bitstream corresponding to the sample to the receiving end, and the receiving end obtains the reconstructed semantic vector corresponding to the sample based on the reference bitstream; An image generation module, configured to generate a pseudo reference image based on the reconstructed semantic vector and the distorted image; A vector acquisition module is used to perform cross-image difference alignment on the pseudo reference image and the distorted image in a preset salient region to obtain a difference feature tensor; The teacher network module is used to obtain inputs to the teacher network's quality regression sub-head and distortion classification sub-head based on the difference feature tensor. The teacher network uses the difference feature tensor to obtain a quality prediction confidence index and a distortion classification prediction confidence index for each image sample in the image sample set, thereby determining the type of each image sample; the types are categorized as low uncertainty, medium uncertainty, and high uncertainty. The student network module is used to freeze the teacher network parameters based on the difference feature tensor and the type of each sample to obtain a trained student network; obtain the difference feature tensor of the test data, use the trained student network to output the subjective quality score and preset several types of distortion probability vectors ,Pick is the dominant distortion type of the test data.
[0017] It can be seen from the above technical solutions that this application has the following advantages: 1. Reduced reference data bandwidth usage: This technology compresses image semantic vectors into a low-bitrate reference stream through product quantization. Compared to traditional full-reference methods that transmit hundreds of KB to MB of lossless images, this technology can compress the reference data to less than 1 / 100 of the original data volume. This compression mechanism enables real-time transmission of reference information within the limited network bandwidth of smart IoT, breaking the bottleneck of IoT link carrying capacity.
[0018] 2. Improved the evaluation accuracy of multi-distortion mixed scenes: By generating pseudo-reference images and aligning cross-image differences, the system effectively captures mixed distortion features such as compression artifacts, random noise, motion blur, and color drift. The teacher network's dual-head architecture (quality regression and distortion classification) combined with an uncertainty assessment mechanism enables the system to maintain high prediction accuracy even in complex distorted scenarios, overcoming the limited accuracy of traditional no-reference methods.
[0019] 3. Real-time computing of edge devices is realized: By adopting a self-distillation training strategy and freezing the parameters of the teacher network to train a lightweight student network, the number of model parameters can be controlled below 5M. Combined with the efficient computational characteristics of product quantization, this application can achieve real-time evaluation on MCUs or edge gateways with limited computing power, meeting the strict real-time requirements of smart IoT terminals. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] In order to more clearly illustrate the technical solution of the present invention, the following is a brief introduction to the drawings required for the description. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0021] Figure 1 This is a flow chart of a low-bitrate semi-reference image quality inspection method based on self-distillation provided in an embodiment of the present application.
[0022] Figure 2 This is a schematic diagram of the internal structure of a low-bitrate semi-reference image quality inspection system based on self-distillation provided in an embodiment of the present application. DETAILED DESCRIPTION
[0023] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.
[0024] It should be understood by those skilled in the art that the embodiments described below are merely preferred embodiments of the present disclosure and do not imply that the present disclosure can only be implemented through these preferred embodiments. These preferred embodiments are merely intended to explain the technical principles of the present disclosure and are not intended to limit the scope of protection of the present disclosure. Based on the preferred embodiments provided by the present disclosure, all other embodiments obtained by those skilled in the art without creative effort should still fall within the scope of protection of the present disclosure.
[0025] It should also be noted that the terms "comprises," "includes," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, method, commodity, or apparatus that includes a series of elements includes not only those elements but also other elements not explicitly listed, or includes elements inherent to such process, method, commodity, or apparatus. In the absence of further limitations, an element defined by the phrase "comprises a ..." does not exclude the presence of other identical elements in the process, method, commodity, or apparatus that includes the element.
[0026] The technical solutions proposed in the embodiments of the present application are described in detail below with reference to the accompanying drawings.
[0027] The embodiment provides a low bit rate semi-reference image quality inspection method based on self-distillation, such as Figure 1 As shown, the method provided in the embodiment of the present application mainly includes the following steps: Step 110: Acquire an intelligent visual terminal image sample set, and perform normalization processing to obtain a normalized image.
[0028] The image sample set is composed of several image samples.
[0029] In some embodiments, this step may specifically include: Image acquisition: The terminal camera captures the reference original image (Added to the image sample set), the resolution is fixed at 1280*720 pixels, that is, the image height is 720 pixels, the width is 1280 pixels, and the frame format is YUV-420.
[0030] Noise suppression: A 3×3 Gaussian filter kernel is applied to each of the three chrominance components to completely remove high-frequency noise.
[0031] Contrast enhancement: Apply histogram equalization to adjust the brightness distribution to make the pixel grayscale uniform.
[0032] Format unification: Convert the processed image to RGB and normalize it to the range [-1, 1] according to the following formula.
[0033] Step 120: Use the frozen basic image encoder network model to extract the image semantic vector of the normalized image.
[0034] In some embodiments, this step may specifically include: The public CLIP-ViT-Base model is used as the basic image encoder, and its parameters are frozen on the terminal side. 512-dimensional semantic vector. This is because, compared to traditional edge or histogram features, semantic features can simultaneously encode scene categories, color layout, and spatial structure, providing high-level priors for semi-reference evaluation. In addition, freezing the network avoids terminal-side retraining and maintains cross-scenario robustness. The specific process is as follows: Network input: Feed into the base image encoder.
[0035] Semantic vector output: Read the CLS Token of the last layer of the basic image encoder network to obtain the semantic vector. The formula is as follows: .
[0036] in, is the frozen basic image encoder network model; Represents the image semantic vector; Represents the i-th semantic dimension (real number).
[0037] Vector normalization: Perform L2 normalization to ensure that the subsequent quantization error is minimized.
[0038] Step 130: Perform product quantization on the image semantic vector to obtain a reference bitstream.
[0039] It should be noted that this step aims to convert the image semantic feature vector into Compressed into a compact bit stream, at the same time, in order to maintain the semantic consistency between the vectors during the compression process, the cosine similarity between the reconstructed vector and the original image semantic vector is required to be no less than 0.94. The specific method is to use two-level product quantization (residual product quantization, denoted as rPQ) combined with entropy coding. First, Performing a first-level quantization, in some embodiments, this step may specifically include: (1) Sub-vector division: 512-dimensional image semantic vector Divide into m=16 32-dimensional sub-vectors in sequence The specific division formula is as follows: ; (2) Index determination: for each sub-vector In the corresponding codebook Calculate the minimum Euclidean distance match and get the index: ; Where K represents the number of codewords in each codebook; =[ ] , represents the mth codebook; {1,...,256} represents the optimal codeword first-level quantization index.
[0040] (3) Bitstream encapsulation: save the 16 indexes in sequence as a 16-byte unsigned integer array, First-level quantization reconstruction vector Most of the semantic information is compressed, but there are still residual errors To eliminate this error, we introduce the second-level residual quantization: we continue to write the residual as , and configure a small codebook for each sub-residual , the small codebook is set to 16 dimensions in this method. The following operations are performed on each sub-residual: ; is the secondary residual index, which indicates the codeword index matched by the mth residual subvector in the secondary codebook. represents the first-level quantized residual of the m-th sub-vector, represents the jm-th codeword (32-dimensional sub-residual vector) in the m-th residual codebook.
[0041] Then concatenate the 16 sub-residual vectors to obtain the secondary quantized residual vector : ; represents the secondary quantized residual vector, ; To improve compression efficiency, the system quantizes the first-level index of each sub-vector and is the secondary residual index The index sequence is then compressed into a reference bitstream b through RangeCoder entropy coding. When the receiving end recovers, the index sequence is decoded to obtain the reconstructed semantic vector .
[0042] Therefore, the two-level quantization consumes M(log2256+log216)=16(8+4)=192bit indexes in total; after RangeCoder entropy coding, the average bit rate is compressed to 140-180Byte, and cos( ,z)≥0.94.
[0043] This scheme ends with two-level quantization, and then directly performs entropy coding on all index sequences to maximize redundancy compression and output a reference bit stream that meets the dual constraints of bandwidth and accuracy.
[0044] Step 140: The distorted image corresponding to the image sample and the reference bitstream are synchronously sent to a receiving end, and the receiving end obtains a reconstructed semantic vector corresponding to the image sample according to the reference bitstream.
[0045] In some embodiments, this step may be specifically as follows: In this step, the receiving end synchronously receives the distorted image and the compressed semantic reference stream, performs a decoding operation and recovers the reconstructed semantic vector The receiver needs to obtain both the distorted image (the image after compression or channel noise) and the reference bitstream b simultaneously to prepare for the subsequent generation of the pseudo reference. The decoding process must be completely symmetrical with the encoding process to ensure no information is lost. The specific process is as follows: Image reception: The gateway receives distorted images through the MQTT protocol.
[0046] Bitstream parsing: In the S3 module, the image semantic vector of the original image is compressed into a reference bitstream b. In this module, this reference bitstream b will be decoded into a reconstructed semantic vector : .
[0047] Step 150: Generate a pseudo reference image based on the reconstructed semantic vector and the distorted image.
[0048] It should be noted that, when only the semantic vector is reconstructed Under these conditions, a “pseudo-reference” with the same structure as the reference image can be quickly reconstructed through a few-step inverse diffusion. In this way, a comparison benchmark can be obtained without transmitting the original image. Therefore, this step is based on reconstructing the semantic vector Fast generation of pseudo reference images.
[0049] In some embodiments, this step may specifically include: Text prompt mapping: Input mapping network , output text prompt t.
[0050] Diffusion initialization: As the starting noise, load the Latent Diffusion model.
[0051] Reverse diffusion: Perform 5 steps of DPM-Solver sampling to obtain a pseudo reference image .in, Represents the DPM-Solver reverse sampler in the fifth step of the Latent Diffusion model. ∈ represents the pseudo reference image.
[0052] Step 160: Perform cross-image difference alignment on the pseudo reference image and the distorted image within a preset salient region to obtain a difference feature tensor.
[0053] In some embodiments, this step performs cross-image difference alignment between the pseudo-reference image and the distorted image within a pre-set salient region. The goal is to precisely compare the structural and color differences between the pseudo-reference image and the distorted image within the salient region and output high-dimensional difference features for use by the scoring network. A sparse strategy is used to significantly reduce the computation of 80% of unimportant patches. The specific process can be as follows: Patch embedding: Divide the two images into 16×16 patches and obtain the token sequence as follows: ; in, Represents the Token sequence obtained after the distorted image is segmented and embedded in the image. Represents a pseudo reference image The token sequence obtained after the same processing, N represents the number of patches into which the image is cut, d is the embedding dimension of each patch, and d is set to 512 in this method.
[0054] Salient selection: Calculating the token sequence of the distorted image Self-attention weight , and select the Top-20%Patch to form the significant set β.
[0055] ; in, represents the average attention of the i-th patch; , represents the h-th head query / key weight matrix; in addition, according to The index of the top-20% patch is selected to form the significant set β.
[0056] Difference coding: According to the significant set β, the patch sequence in the significant region is Input Cross-ViT for cross-image alignment and fusion encoding, and output the difference feature tensor D, which represents the difference features between the pseudo reference image and the distorted image in the salient area.
[0057] .
[0058] Step 170: Based on the difference feature tensor, the input of the teacher network quality regression sub-head and the distortion classification sub-head is obtained, and using the difference feature tensor, the teacher network obtains the quality prediction confidence index and the distortion classification prediction confidence index of each image sample in the image sample set, and then determines the type of each image sample.
[0059] This step uses a lightweight student network to simultaneously output subjective quality scores and distortion types (blur, noise, compression artifacts, color distortion, geometric deformation, contrast distortion, brightness distortion, and edge blur). During training, it is distilled by a large teacher network, and during inference, only the student network is run to adapt to the low-power processor, and the final output is the quality score and distortion type. However, the intelligent vision terminal has few samples and large quality fluctuations. Therefore, the teacher network's predictions for some samples are unreliable (such as when the pseudo-reference image is blurred), and traditional equal distillation will amplify the noise. This method proposes a high-variance trigger protection mechanism for the teacher network to prevent students from learning incorrect knowledge. Therefore, this step also introduces an uncertainty weighting mechanism. Monte Carlo Dropout is used to estimate the teacher network's prediction confidence in the quality score and distortion type, and dynamically adjust the knowledge distillation weights. The specific process is as follows: Pooling operation: perform mean pooling and maximum pooling on the difference feature tensor D to obtain The two serve as the input of the quality regression sub-head and the distortion classification sub-head respectively.
[0060] Teacher network uncertainty estimation: To prevent the teacher network from misleading when the sample quality fluctuates greatly or the pseudo-reference image is severely distorted, this step introduces a prediction uncertainty estimation mechanism based on Dropout to dynamically adjust the distillation weights. Specifically, for the same difference feature input tensor D, Dropout is enabled in the teacher network and M = 3 forward propagations are performed to simulate the prediction volatility of the network output, obtaining three sets of quality prediction results {Q(1), Q(2), Q(3)} and the corresponding distortion type classification probability vectors {c(1), c(2), c(3)}. Among them, for quality prediction, the variance of the quality prediction scores is calculated three times to obtain The smaller the variance, the more stable the model judgment and the higher the reliability. This value is used as a quality prediction confidence indicator and is involved in subsequent distillation weight adjustment.
[0061] Among them, for the distortion type classification, the entropy value of the classification probability vector is calculated three times The specific formula is as follows: ; It is used as a confidence indicator for distortion classification prediction.
[0062] Calculation of uncertainty thresholds for quality prediction and distortion classification: To determine the uncertainty thresholds for image quality prediction and distortion classification, this step selects image samples collected from intelligent visual terminals (such as cameras, inspection cameras, agricultural cameras, etc.) to construct a validation set. The specific steps are as follows: Validation set sample preprocessing: The validation set images go through the following steps in sequence: After S110-S160 processing, the difference feature tensor D corresponding to each image is finally obtained; the difference feature tensor D of each sample will be used as the input of the teacher network for uncertainty assessment.
[0063] The teacher network propagates forward multiple times and collects prediction results: Repeat the forward propagation M=3 times for the feature tensor D of each image, turning on Dropout each time, and record the quality prediction score sequence {Q(1), Q(2), Q(3)} and the distortion classification probability distribution sequence {c(1), c(2), c(3)}.
[0064] Calculate the variance of the quality prediction score and the distortion classification prediction entropy of each sample. Furthermore, for all image samples in the validation set, collect the following: Constitute a set of distributions, classification prediction entropy This constitutes another set of distributions.
[0065] Extract the high percentile of the distribution as the uncertainty threshold: By counting the 95% percentiles of the above two distributions, the following threshold is defined: Quality prediction volatility threshold , classification prediction confidence threshold .
[0066] In this method, the quality prediction confidence index ≤ And the distortion classification prediction confidence index ≤ , the sample is classified as low uncertainty; if only one item exceeds the threshold, it is classified as medium uncertainty; if both items exceed the threshold significantly at the same time, it is recorded as high uncertainty.
[0067] Step 180: Based on the difference feature tensor and the type of each sample, freeze the teacher network parameters to obtain the trained student network; obtain the difference feature tensor of the test data, use the trained student network to output the subjective quality score and several preset distortion probability vectors ,Pick is the dominant distortion type of the test data.
[0068] It should be noted that the student network performs independent predictions. At this point, the inference end no longer relies on the teacher network, and all operations are completed within the student network. The process is as follows: The student network performs the forward propagation process: First, the input D is still subjected to mean pooling and maximum pooling, and the result is: ; in, () is the mean pooling function, () is the maximum pooling function. Next, the quality sub-head performs two layers of full connection, and the dimension changes from 384 to 64 to 1, and outputs the subjective quality prediction as follows: ; in, = represents a lightweight fully connected sub-head, i.e., a two-layer fully connected network forward propagation model function. Then, the classification sub-head performs two layers of full connection, and the dimension changes from 384 to 64 to 8. The distortion probability is obtained by Softmax, as follows: ; Uncertainty Weighted Distillation Loss: ; in, 、 Output quality scores and classification results for students; 、 The quality score of the teacher network output and the classification result are output by the teacher; KL represents the Kullback-Leibler divergence. Through this step, the student network and codebook can be updated synchronously, while the teacher network parameters remain frozen.
[0069] in , The weights are set as follows: ; ; Among them, when the sample falls into the "high uncertainty" state, the distillation weight is set to zero, and the gradient update depends only on the student's own supervision signal; when it is in the "medium uncertainty" state, the weight is set to zero. and The system gradually decreases the error, naturally reducing the impact of samples with large errors on the student network. Low-uncertainty samples remain fully distilled to fully utilize the teacher information. If excessive high-uncertainty samples appear for k consecutive epochs, the system temporarily freezes the corresponding sub-head parameters and automatically unfreezes them after the mixed distortion scenario is alleviated, preventing the model from being biased by noisy gradients.
[0070] Furthermore, during the training process, each time the loss function (such as distillation loss, quality score loss, and classification loss) is calculated, the learning rate and weight update are automatically adjusted based on the model's feedback on the loss.
[0071] Parameter update: Backpropagation only updates the student network and the quantization codebook; the teacher network is always frozen.
[0072] Student network independent reasoning stage: pseudo reference generation: terminal local reconstruction , single-step diffusion gives ; Next is the difference feature extraction: run sparse HAX-ViT (6 layers, Top 20% Token) to make student predictions, and finally output the subjective quality score (0-100), and eight categories (i) of distortion probability vectors ,Pick is the dominant distortion type. Represents the predicted probability of the i-th distortion type, which is output from the student or teacher network through the Softmax layer.
[0073] This application limits the "teacher network confidence gated distillation" to the training phase, and allows the student network to independently and deterministically perform the above-mentioned mean, maximum pooling + lightweight fully connected reasoning during the inference phase. This module not only avoids mis-distillation when the teacher output is unreliable, but also ensures the real-time and stability of terminal deployment, forming a complete closed loop.
[0074] Through the above steps, the embodiment of the present application achieves high-precision quality assessment of intelligent visual terminal images using only a low-bitrate reference bit stream. It has the technical advantages of extremely low reference information volume, low computational overhead, and robustness to multiple types of distortion. It can realize real-time quality monitoring under narrowband links such as NB-IoT, LoRa, and cellular Cat-M.
[0075] In addition, this application Figure 2 The embodiment of the present application provides a low-bitrate semi-reference image quality inspection system based on self-distillation. Figure 2 As shown, the system provided in the embodiment of the present application mainly includes: An image acquisition module 210 is used to acquire a sample set of an intelligent visual terminal and perform a normalization process to obtain a normalized image; A vector extraction module 220 is configured to extract a normalized image semantic vector using a frozen basic image encoder network model; The code stream obtaining module 230 is used to perform product quantization on the image semantic vector to obtain a reference code stream; The vector reconstruction module 240 is used to synchronously send the distorted image corresponding to the sample and the reference bitstream to the receiving end, and the receiving end obtains the reconstructed semantic vector corresponding to the sample based on the reference bitstream; An image generation module 250 is configured to generate a pseudo reference image based on the reconstructed semantic vector and the distorted image; A vector acquisition module 260 is configured to perform cross-image difference alignment on the pseudo reference image and the distorted image within a preset salient region to obtain a difference feature tensor; A teacher network module 270 is configured to obtain inputs to a teacher network quality regression sub-head and a distortion classification sub-head based on the difference feature tensor, and to use the difference feature tensor to obtain a quality prediction confidence index and a distortion classification prediction confidence index for each image sample in the image sample set, thereby determining a type of each image sample; wherein the types are classified into low uncertainty, medium uncertainty, and high uncertainty; The student network module 280 is used to freeze the teacher network parameters based on the difference feature tensor and the type of each sample to obtain a trained student network; obtain the difference feature tensor of the test data, and use the trained student network to output the subjective quality score and several preset distortion probability vectors. ,Pick is the dominant distortion type of the test data.
[0076] In addition, an embodiment of the present application also provides a non-volatile computer storage medium on which executable instructions are stored. When the executable instructions are executed, a low-bitrate semi-reference image quality inspection method based on self-distillation as described above is implemented.
[0077] The above description of the disclosed embodiments is intended to enable one skilled in the art to implement or use the present invention. Various modifications to these embodiments will be readily apparent to one skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention is not limited to the embodiments shown herein but is intended to conform to the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A low-bitrate semi-reference image quality inspection method based on self-distillation, characterized in that: The method comprises: Acquire an image sample set of an intelligent visual terminal and perform a normalization process to obtain a normalized image; wherein the image sample set is composed of a plurality of image samples; A frozen basic image encoder network model is used to extract the image semantic vector of the normalized image; the image semantic vector is product-quantized to obtain a reference bitstream; The distorted image corresponding to the image sample and the reference bitstream are synchronously sent to the receiving end, and the receiving end obtains the reconstructed semantic vector corresponding to the image sample based on the reference bitstream; based on the reconstructed semantic vector and the distorted image, a pseudo reference image is generated; Perform cross-image difference alignment on the pseudo reference image and the distorted image in a preset salient region to obtain a difference feature tensor; Based on the difference feature tensor, the teacher network obtains inputs to the quality regression sub-head and the distortion classification sub-head. Using the difference feature tensor, the teacher network obtains a quality prediction confidence index and a distortion classification prediction confidence index for each image sample in the image sample set, thereby determining the type of each image sample. The types are categorized as low uncertainty, medium uncertainty, and high uncertainty. Based on the difference feature tensor and the type of each sample, the teacher network parameters are frozen to obtain a trained student network. Obtain the difference feature tensor of the test data, use the trained student network to output the subjective quality score and preset several types of distortion probability vectors ,Pick is the dominant distortion type of the test data.
2. The low-bitrate semi-reference image quality inspection method based on self-distillation according to claim 1, characterized in that: Obtain a sample set of the intelligent visual terminal and perform standardization processing to obtain a normalized image, including: Terminal camera collects intelligent visual terminal sample set , the resolution is fixed at 1280*720 pixels and the frame format is YUV-420; right Apply 3×3 Gaussian filter kernel to the three chroma components respectively; Apply histogram equalization to adjust the brightness distribution to make the pixel grayscale uniform; Convert the processed image to RGB and normalize it to the range [-1, 1] according to the following formula.
3. The low-bitrate semi-reference image quality inspection method based on self-distillation according to claim 1, characterized in that: Using the frozen basic image encoder network model, the image semantic vector of the normalized image is extracted, specifically including: The normalized image Feed into the basic image encoder; through the formula: , read the last layer CLS Token of the basic image encoder network to obtain the image semantic vector ; in, is the frozen basic image encoder network model; Represents the i-th semantic dimension.
4. The low-bitrate semi-reference image quality inspection method based on self-distillation according to claim 1, characterized in that: The image semantic vector is product-quantized to obtain a reference bitstream, which includes: According to the preset division formula: , divide the image semantic vector into m=16 32-dimensional sub-vectors in turn ; According to the preset formula: , for each In the corresponding codebook Calculate the Euclidean distance and obtain the first-level quantitative index ; Where K represents the number of codewords in each codebook; =[ ] , represents the mth codebook; {1,...,256} represents the first-level quantization index; Save the 16 first-level quantization indexes in sequence as a 16-Byte unsigned integer array to obtain the first-level quantization reconstruction vector ; Perform the operation on each sub-residual: ;in, represents the secondary residual index, represents the first-level quantized residual of the m-th sub-vector, represents the jm-th codeword in the m-th residual codebook; The first-level quantization index of each sub-vector and is the secondary residual index The index pairs are combined into index pairs, and the index pairs are compressed into reference bitstream b through RangeCoder entropy coding.
5. The low-bitrate semi-reference image quality inspection method based on self-distillation according to claim 1, characterized in that: The receiving end obtains the reconstructed semantic vector based on the reference bitstream, specifically including: The receiving end obtains reference stream b: By formula: , decode the reference bitstream into a reconstructed semantic vector ; Represents the decoding function.
6. The low-bitrate semi-reference image quality inspection method based on self-distillation according to claim 1, characterized in that: Generate a pseudo reference image based on the reconstructed semantic vector and the distorted image, specifically including: The semantic vector will be reconstructed Input mapping network , output text prompt t; Distorted image As the starting noise, load the Latent Diffusion model; Perform 5 steps of DPM-Solver sampling to obtain a pseudo reference image ;in, Represents the 5th step DPM-Solver reverse sampler of the Latent Diffusion model, ∈ represents the pseudo reference image.
7. The low-bitrate semi-reference image quality inspection method based on self-distillation according to claim 1, characterized in that: Perform cross-image difference alignment on the pseudo reference image and the distorted image in the preset salient region to obtain the difference feature tensor, specifically including: The pseudo reference image and the distorted image are divided into 16×16 patches, and the token sequence is obtained as follows: ; By formula: , calculate the self-attention weight ; And select the top-20% patches to form the set β; in, represents the average attention of the i-th patch; , represents the h-th head query / key weight matrix; Input 6 layers of Cross-ViT to the token in β and output the difference feature tensor D.
8. The low-bitrate semi-reference image quality inspection method based on self-distillation according to claim 1, characterized in that: Using the difference feature tensor, the teacher network obtains the quality prediction confidence index and distortion classification prediction confidence index of each sample in the sample set, and then determines the type of each sample, including: For the same difference feature tensor D, Dropout is enabled in the teacher network and M = 3 forward propagations are performed to simulate the prediction volatility of the network output, obtaining 3 sets of quality prediction results {Q(1), Q(2), Q(3)} and the corresponding distortion type classification probability vectors {c(1), c(2), c(3)}; Calculate the variance of the quality prediction scores three times and get , as a quality prediction confidence indicator; By formula: , calculate the entropy value of the three classification probability vectors It will be used as a confidence indicator for the distortion classification prediction; According to the quality prediction confidence index and the distortion classification prediction confidence index of all samples, the quality prediction confidence index distribution and the distortion classification prediction confidence index distribution are obtained; According to the preset quantiles of the distribution, the quality prediction volatility threshold and the classification prediction confidence threshold are obtained; When the quality prediction confidence index of a sample is less than or equal to the quality prediction volatility threshold, and the distortion classification prediction confidence index is less than or equal to the classification prediction confidence threshold, the sample is of low uncertainty type; When the quality prediction confidence index of a sample is greater than the quality prediction volatility threshold or the distortion classification prediction confidence index is greater than the classification prediction confidence threshold, the sample is of medium uncertainty type; When the quality prediction confidence index of a sample is greater than the quality prediction volatility threshold and the distortion classification prediction confidence index is greater than the classification prediction confidence threshold, the sample is of high uncertainty type.
9. The low-bitrate semi-reference image quality inspection method based on self-distillation according to claim 1, characterized in that: Based on the difference feature tensor and the type of each sample, the student network is used to output the subjective quality score and several preset distortion probability vectors. ,Pick The dominant distortion types include: Input difference feature tensor, after mean pooling and maximum pooling, we get: ;in, () is the mean pooling function, () is the maximum pooling function; based on , perform two-layer full connection on the quality sub-head, dimension change: 384 → 64 → 1, output subjective quality prediction ; based on , the classification sub-head performs two-layer full connection, the dimension changes to: 384 → 64 → 8, and the distortion probability is obtained by Softmax ;in, ; According to the formula: , calculate the uncertainty weighted distillation loss; in, 、 The subjective quality prediction and distortion probability of the teacher network output; in , The weights are set as follows: , , in, Indicates the preset adjustment factor; Backpropagation only updates the student network and the quantized codebook; the teacher network is always frozen; Access to a network of trained students; Input the difference feature tensor into the trained student network and output the subjective quality score and distortion probability vector ,Pick is the dominant distortion type.
10. A low-bitrate semi-reference image quality inspection system based on self-distillation, characterized in that: The system comprises: An image acquisition module is used to acquire a sample set of an intelligent visual terminal and perform standardization processing to obtain a normalized image; A vector extraction module is used to extract normalized image semantic vectors using a frozen basic image encoder network model; A code stream acquisition module is used to perform product quantization on the image semantic vector to obtain a reference code stream; The vector reconstruction module is used to synchronously send the distorted image and reference bitstream corresponding to the sample to the receiving end, and the receiving end obtains the reconstructed semantic vector corresponding to the sample based on the reference bitstream; An image generation module, configured to generate a pseudo reference image based on the reconstructed semantic vector and the distorted image; A vector acquisition module is used to perform cross-image difference alignment on the pseudo reference image and the distorted image in a preset salient region to obtain a difference feature tensor; The teacher network module is used to obtain inputs to the teacher network's quality regression sub-head and distortion classification sub-head based on the difference feature tensor. The teacher network uses the difference feature tensor to obtain a quality prediction confidence index and a distortion classification prediction confidence index for each image sample in the image sample set, thereby determining the type of each image sample; the types are categorized as low uncertainty, medium uncertainty, and high uncertainty. The student network module is used to freeze the teacher network parameters based on the difference feature tensor and the type of each sample to obtain a trained student network; obtain the difference feature tensor of the test data, use the trained student network to output the subjective quality score and preset several types of distortion probability vectors ,Pick is the dominant distortion type of the test data.
Citation Information
Patent Citations
Semantic and structural distillation reference-free image quality evaluation method
CN117274173A
Domain distribution self-adaption-oriented reference-free image quality evaluation method and system
CN117745640A
Interactive semantic-aware self-learning framework and interpretable visual recognition method
WO2025091768A1