Image quality evaluation method, electronic device, storage medium, and program product
Patent Information
- Application Number
- CN202510759089.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-09
- Publication Date
- 2026-08-28
- Estimated Expiration
- 2045-06-09
AI Technical Summary
[0005]本申请实施例提供一种图像质量评价方法、电子设备、存储介质及程序产品,能够解决当前视频图像质量评价精度、实时性方面有缺陷,自动化程度不高,也不够智能化,导致视频图像的质量评价效率较低的问题
[0010]第五方面,本申请实施例提供了一种计算机程序产品,所述计算机程序产品包括存储在非暂态计算机可读存储介质上的计算机程序,所述计算机程序包括程序指令,当所述程序指令被计算机执行时,使所述计算机执行如第一方面所述的方法的步骤。
Smart Images

Figure CN121120478B_ABST
Abstract
Description
Technical Field
[0001] This application belongs to the field of image processing technology, specifically relating to an image quality evaluation method, electronic device, storage medium, and program product. Background Technology
[0002] In some areas, traffic complexity is gradually increasing, and traffic monitoring systems are becoming more sophisticated. Surveillance video has become a crucial basis for relevant departments to mediate traffic disputes, determine violations, and plan urban roads. Tens of thousands of traffic monitoring cameras record and store massive amounts of video data daily. Relevant departments hope to establish a reasonable video image quality assessment system, focusing on key indicators such as image quality, while keeping costs under control. This will ensure the usability of video quality and avoid losses and wasted investment caused by video image failure.
[0003] Currently, most mainstream video image quality assessment methods employ traditional image processing techniques, primarily processing image data and evaluating its quality on the backend. This is because traditional image processing techniques are relatively mature; they can be directly applied to different scenarios with minor modifications to existing application programming interfaces (APIs), at a low cost, and offer fast speeds for scenarios where accuracy requirements are not high. Secondly, the non-embedded modules used in the backend systems have ample memory to meet the space requirements of more modules, run more algorithms, and are easier to expand in terms of hardware and software, offering good compatibility and allowing for the replacement of appropriate hardware based on algorithm requirements.
[0004] Traditional image processing technology is based on digital image processing. The main method for identifying video image quality problems is to use a large amount of labeled data for supervised training. This process requires frequent manual screening and annotation of video images on a large scale, wasting a lot of time and manpower. It also has shortcomings in terms of accuracy and real-time performance, and its automation and intelligence are not high enough, resulting in low efficiency in video image quality assessment. Summary of the Invention
[0005] This application provides an image quality evaluation method, electronic device, storage medium, and program product, which can solve the problems of current video image quality evaluation having deficiencies in accuracy and real-time performance, low automation, and insufficient intelligence, resulting in low efficiency in video image quality evaluation.
[0006] In a first aspect, embodiments of this application provide an image quality evaluation method, which includes: performing a distortion transformation on a target image until the distorted image obtained by the distortion transformation of the target image meets preset conditions; inputting the target image and the distorted image into an image quality evaluation network to obtain the quality evaluation scores and distortion types of the target image and the distorted image.
[0007] Secondly, embodiments of this application provide an image quality evaluation device, which includes: a transformation module for performing a distortion transformation on a target image until the distorted image obtained by the distortion transformation of the target image meets preset conditions; and an evaluation module for inputting the target image and the distorted image into an image quality evaluation network to obtain the quality evaluation scores and distortion types of the target image and the distorted image.
[0008] Thirdly, embodiments of this application provide an electronic device including a processor, a memory, and a program or instructions stored in the memory and executable on the processor, wherein the program or instructions, when executed by the processor, implement the steps of the method described in the first aspect.
[0009] Fourthly, embodiments of this application provide a readable storage medium on which a program or instructions are stored, which, when executed by a processor, implement the steps of the method described in the first aspect.
[0010] Fifthly, embodiments of this application provide a computer program product, the computer program product including a computer program stored on a non-transitory computer-readable storage medium, the computer program including program instructions, which, when executed by a computer, cause the computer to perform the steps of the method described in the first aspect.
[0011] In this embodiment, the target image is distorted until the distorted image obtained by the distortion transformation meets preset conditions. The target image and the distorted image are input into an image quality evaluation network to obtain the quality evaluation scores and distortion types of the target image and the distorted image. Through a feedback control mechanism, the quality of the generated distorted image is controlled, and the distorted image is generated efficiently and with high quality. There is no need for manual annotation of image quality data. The image quality evaluation network automatically outputs the quality evaluation score and distortion type of the target image, thereby improving the efficiency of image quality evaluation. Attached Figure Description
[0012] Figure 1 This is a flowchart illustrating an image evaluation method provided in an embodiment of this application; Figure 2 This is a schematic diagram of a process for generating a distorted image provided in an embodiment of this application; Figure 3 This is a schematic diagram of a feedback mechanism provided in an embodiment of this application; Figure 4 This is a diagram of an image quality evaluation network structure provided in an embodiment of this application; Figure 5This is a schematic diagram of a Conformer structure provided in an embodiment of this application; Figure 6 This is a schematic diagram of the structure of an image evaluation device provided in an embodiment of this application; Figure 7 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Detailed Implementation
[0013] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0014] The terms "first," "second," etc., used in the specification and claims of this application are used to distinguish similar objects and not to describe a specific order or sequence. It should be understood that such use of data can be interchanged where appropriate so that embodiments of this application can be implemented in orders other than those illustrated or described herein, and the objects distinguished by "first," "second," etc., are generally of the same class and the number of objects is not limited; for example, a first object can be one or more. Furthermore, in the specification and claims, "and / or" indicates at least one of the connected objects, and the character " / " generally indicates that the preceding and following objects are in an "or" relationship.
[0015] Figure 1 This application illustrates an embodiment of an image quality evaluation method provided. This method can be executed by an electronic device, which may include a server and / or a terminal device. In other words, the method can be executed by software or hardware installed on the electronic device, and includes the following steps: Step S102: Perform a distortion transformation on the target image until the distorted image obtained by the distortion transformation of the target image meets the preset conditions.
[0016] The image quality evaluation method provided in this application can be applied to surveillance videos in traffic monitoring scenarios. Subsequent embodiments will use surveillance videos in traffic monitoring scenarios as an example for illustration. The image quality averaging method provided in this application can also be applied to video images in other scenarios; therefore, the application scenarios of the image quality averaging method in this application are not specifically limited here.
[0017] This application addresses the problem of image quality assessment in real-world traffic monitoring videos. It achieves image quality evaluation and distortion type determination without manual annotation, significantly reducing the difficulty of manual data screening and improving efficiency. To obtain a large amount of near-real traffic image data, a Feedback Mechanism for Image Quality Assessment (FMIQA) algorithm is designed. This algorithm constructs a large dataset of real-world traffic monitoring images, generating images with high quality variations while meeting the data requirements of this experiment.
[0018] The FMIQA algorithm designed in this application constructs the self-supervised dataset used in this application embodiment. By adding a feedback control mechanism and combining it with the SSIM method, the quality of the generated image data is controlled, resulting in a high-quality dataset that is generated efficiently and in a variety of image data types. The key point lies in the clever combination of the feedback mechanism, which is equivalent to image augmentation with automatic control. FMIQA can adaptively adjust distortion-related parameters to control the quality of the generated distorted images, ensuring the quality of the distorted images. Data is crucial to the performance of AI algorithms; through quality control, the dataset can achieve high quality.
[0019] To construct a real-world dataset, the FMIQA algorithm was designed based on the publicly available image quality assessment dataset TID2013 to generate distorted images with different quality conditions, weather conditions, and degrees of distortion in real-world traffic monitoring scenarios. Due to the characteristics of the generation method, the generated distorted images contain two labels: distortion type and distortion degree. This data will be used for self-supervised training of the model.
[0020] In this embodiment, a custom number of video frames can be extracted from, for example, traffic monitoring videos. Images with no signal, completely black and white, or blank images are filtered out, and the remaining images are the target images. In this embodiment, images of different climates and seasons can be generated based on the target images using data augmentation methods, such as adding fog, rain, or snow to the target images to enrich the data types and enable the model to adapt to more diverse road traffic environments. Simultaneously, the FMIQA method is used to control the quality of the generated images, merging the augmented data with the original data. Then, the FMIQA algorithm is used to perform distortion transformation on the target images.
[0021] In one implementation, the distortion transformation includes at least one of the following: brightness transformation, color cast transformation, blur transformation, occlusion transformation, black and white transformation, screen tearing transformation, and noise transformation.
[0022] In this embodiment, seven anomaly types are defined: brightness anomaly, color cast anomaly, blur anomaly, occlusion anomaly, black and white anomaly, screen distortion anomaly, and noise anomaly. Each anomaly type is further divided into five distortion levels. Therefore, this embodiment provides seven distortion transformations: brightness transformation, color cast transformation, blur transformation, occlusion transformation, black and white transformation, screen distortion transformation, and noise transformation. At least one of these seven distortion transformations can be selected to perform a distortion transformation on the target image. This embodiment also allows for the deletion of images that cannot fully generate the five distortion levels. This embodiment uses a noise anomaly (noise transformation) algorithm as an example to demonstrate part of the algorithm's flow.
[0023] For noise transformation, various methods such as additive Gaussian noise, comfortable noise, lossy compression, spatially correlated noise, and high-frequency noise can be used to enrich the types of blur distortion, making the resulting distorted image more adaptable to changing and complex real-world environments. In practical applications, one method can be selected for blur transformation at a time.
[0024] Figure 2 The diagram shows a flowchart for generating a distorted image with comfort noise. Comfort noise is generated by converting the target image into a YCbCr form, and then processing the Y, Cb, and Cr components separately to obtain a blurred, distorted image. Taking the Y channel as an example: the Y channel is compressed using an Adaptive Discrete Cosine Transform (ADCT), then decompressed to obtain Y. By setting appropriate coefficients, Y is replaced by nY1-Y. The Cb and Cr channels are processed similarly to obtain Cbn and Crn. Finally, the three channels are combined and converted into an RGB channel image to obtain the distorted image with abnormal comfort noise.
[0025] For other methods of generating blur distortion, such as additive Gaussian noise, additive Gaussian noise is added to the target image, followed by ADCT encoding and decompression / unblocking. The distortion level is controlled using the noise standard deviation to obtain a distorted image. Another example is spatially correlated noise, which is achieved by altering the RGB color components of the target image and applying motion blur. The degree of blur and the transform value are used to control the degree of distortion, thus obtaining a distorted image.
[0026] For other types of distortion, appropriate distortion algorithms can be designed and combined with FMIQA. Through various types and degrees of distortion transformation, distorted images of different types and degrees of distortion can be obtained and formed into a dataset for subsequent self-supervised learning of the model.
[0027] In this embodiment of the application, the distorted image obtained after distorting the target image may not meet the actual requirements. Therefore, this embodiment of the application cleverly incorporates a feedback mechanism into FMIQA, that is, it is necessary to determine whether the obtained distorted image meets the preset conditions. If the obtained distorted image does not meet the preset conditions, it is necessary to circulate the distorting transformation of the target image until the obtained distorted image meets the preset conditions.
[0028] In one implementation, the step of distorting the target image until the distorted image obtained by distorting the target image meets a preset condition includes: distorting the target image using preset control parameters; if the similarity value between the obtained distorted image and the target image does not meet the preset condition, adjusting the control parameters, and distorting the distorted image according to the adjusted control parameters until the similarity value between the distorted image and the target image meets the preset condition, wherein the preset condition includes that the similarity value between the distorted image and the target image is within a preset similarity range.
[0029] Figure 3 A schematic diagram of the feedback mechanism provided in an embodiment of this application is shown, such as... Figure 3 As shown, the FMIQA algorithm takes the target image as input. Different distortion transformation methods are designed based on different distortion types, and the degree of distortion transformation is controlled by the control parameter K. First, the target image is distorted to obtain a distorted image. Then, the objective quality assessment method SSIM is used as an indicator to measure the degree of distortion between the distorted image and the target image. The similarity (SSIM) value between the two is calculated. If the SSIM value is too large, exceeding the set threshold Tr, the similarity value between the distorted image and the target image is not within the preset similarity range, indicating that the generated distorted image is not significantly different from the target image, which is not conducive to the model's differentiation of different levels of distortion. Therefore, the generated image needs to be deleted, and the control parameter K is adjusted, and the distortion transformation is performed again. If the SSIM value is too small, less than the set threshold T... l If the similarity value between the distorted image and the target image is not within the preset similarity range, it indicates that the generated distorted image is too distorted. Therefore, this generated image needs to be deleted, the control parameter K needs to be readjusted, and the distortion transformation needs to be performed again, and this process is repeated. If the SSIM value is within the preset similarity range (T... l Within Tr), the generated distorted image is considered to meet the preset conditions.
[0030] In practical applications, up to five different quality levels of distorted images can be generated for the same type of distortion, and the maximum loop can be set to 10 times. If, after the loop ends, the SSIM value between the generated distorted image and the original target image still cannot be within the preset similarity range, and an effective distorted image cannot be obtained, then the target image is considered unsuitable for application in this dataset generation method and can be deleted. Through this feedback mechanism, combined with objective quality evaluation methods to control the quality of the generated distorted images, a certain level of difference in image quality between the distorted image and the target image can be achieved, without completely deviating from the target image.
[0031] The similarity score SSIM is often used to represent the degree of similarity between a distorted image and a target image. Its specific calculation method is as follows:
[0032] x Represents the target image. y These represent distorted images. μ x This represents the mean of the target image. μ y This represents the mean of the distorted image. σ² x Represents the variance of the target image. σ² y The variance of the distorted image is represented by... σ xy This represents the covariance between the target image and the distorted image. c 1 and c 2 is a constant used to avoid a denominator of 0. From the formula above, we can see that SSIM mainly compares distorted and target images in terms of structure, brightness, and contrast. Compared to MSE and PSNR, SSIM better reflects the human eye's evaluation of image quality; a higher value indicates greater similarity between the two.
[0033] The FMIQA method provided in this application cleverly combines a feedback mechanism with image augmentation and automatic control, effectively controlling the quality of the generated data. This results in the efficient and high-quality generation of datasets and enriches the variety of data. The designed FMIQA method can generate a large number of effective datasets for self-supervised learning while saving time. This significantly expands the amount of datasets available, yielding a large amount of high-quality data for subsequent model training.
[0034] Step S104: Input the target image and the distorted image into the image quality evaluation network to obtain the quality evaluation score and distortion type of the target image and the distorted image.
[0035] To meet the needs of real-world traffic monitoring scenarios, reduce reliance on labeled images, and assist relevant personnel in identifying abnormal image types, this application proposes a full-reference IQA method. It learns quality scores and transformation types directly using the target image and the generated distorted image as a dataset, without requiring additional labeled data. In this application embodiment, after obtaining the target image and the distorted image, they can be input into an image quality assessment network. The image quality assessment network outputs the quality assessment scores and distortion types of the target image and the distorted image, such as... Figure 4 As shown, the network quality assessment network employs a full-reference IQA method, requiring a target image for application. The network quality assessment network consists of a pair of Siamese networks and two branch networks. The number of transformation iterations replaces the quality score, and weak labels are constructed based on the distortion transformation type. After feature extraction by the Siamese network, the target image and the distorted image pair are respectively processed by the quality assessment branch and the distortion transformation classification branch to achieve image quality assessment and anomaly type classification functions.
[0036] The image quality assessment method provided in this application involves distorting a target image until the resulting distorted image meets preset conditions. The target image and the distorted image are then input into an image quality assessment network to obtain their quality assessment scores and distortion types. This method cleverly combines a feedback mechanism with image augmentation and automatic control, effectively controlling the quality of the generated data. This results in efficient and high-quality dataset generation and enriches the data types. The designed FMIQA method can generate a large number of effective datasets for self-supervised learning while saving time. This significantly expands the dataset size, yielding a large amount of high-quality data for subsequent model training. It eliminates the need for manual annotation of image quality data, automatically outputting the target image's quality assessment score and distortion type through the image quality assessment network, thus improving image quality assessment efficiency. This application embodiment eliminates the need for manual annotation of image data, removing errors caused by subjective judgment and improving algorithm accuracy. It better overcomes the shortcomings of traditional digital image processing technologies and better aligns with the trend of intelligent transportation.
[0037] In one implementation, the image quality assessment network includes a quality assessment branch network. The step of inputting the target image and the distorted image into the image quality assessment network to obtain quality assessment scores and distortion types for the target image and the distorted image includes: inputting the target image and the distorted image into the image quality assessment network; extracting a first feature of the target image and a second feature of the distorted image; calculating the feature difference between the first feature and the second feature; inputting the feature difference into the quality assessment branch network; and outputting the quality assessment scores for the target image and the distorted image.
[0038] In this embodiment of the application, the target image X can be... ref And its distorted image X after distortion transformation dist As a set of image pairs, input into the Siamese network, the first feature Z of the target image is obtained after feature extraction. ref The second feature Z of the distorted image dist And calculate the first feature Z ref Second feature Z dist The L1 distance between the two features, i.e., the feature difference |Z| between the first and second features. ref -Z dist The feature difference is used as the basis for the quality evaluation branch network, which outputs the quality evaluation scores of the target image and the distorted image.
[0039] The Siamese network uses Inception-ResNet-v2 pre-trained from ImageNet to extract features from both the target and distorted images. It also connects some intermediate feature maps from the Inception-ResNet-v2 layers. After passing through the Siamese network, the target and distorted images respectively obtain Z... ref and Z dist To obtain the difference information between the target image and the distorted image, |Z| is used. ref -Z dist | is used for calculation, and it is used as the input for the subsequent two branches.
[0040] In one implementation, the step of inputting the feature difference into the quality assessment branch network and outputting quality assessment scores for the target image and the distorted image includes: transforming the feature difference through the decoder-encoder module of the quality assessment branch network to obtain a difference image; inputting the difference image into a multi-binary classifier to determine a reference quality score for the difference image; determining a prediction probability that the score threshold of each binary classifier is less than or equal to the reference quality score based on the reference quality score and the score threshold of each binary classifier in the multi-binary classifier; and determining the average of the prediction probabilities of each binary classifier as the quality assessment score.
[0041] We can first consider the feature difference |Z ref -Z dist The input is first fed into the encoder-decoder module of the quality assessment branch network. The encoder-decoder module transforms the feature differences to obtain the difference image. Specifically, the encoder-decoder uses a single ConformerBlock. The Conformer network adds a CNN structure to the Transformer network, combining the local information association capabilities of CNNs with the global information interaction capabilities of Transformers to achieve efficient utilization of both local and global features. The overall structure of the ConformerEncoder is as follows: Figure 5 As shown, its main component is the ConformerBlock, which, compared to the TransformerEncoder, adds an extra feedforward network and a convolutional module. The entire structure consists of a first feedforward network, a multi-head attention module, a convolutional module, and a second feedforward network. Input x i When passing through ConformerBlock, the calculation is as follows:
[0042] Special differences |Z can be used ref -Z dist After the input is fed into the Encoder-Decoder module to obtain the difference image, it is then passed through a Multiple Binary Classifier (MBC) to obtain the similarity index of the input image pair. The similarity index obtained from this branch is matched with the label value to learn the image quality score. The loss function is BCELoss.
[0043] Specifically, the difference images are input into a multiple binary classifier (MBC), which predicts the quality score. Instead of a traditional regressor, MBC uses N binary classifiers to learn the quality score of the generated image. Specifically, the classifiers are trained based on whether the reference quality score of the generated difference image is greater than the score threshold of each binary classifier in the multiple binary classifier.
[0044] After inputting the difference image into a multi-binary classifier to determine the reference quality score of the difference image, the method further includes: determining a score threshold for each binary classifier based on the number of each binary classifier in the multi-binary classifier.
[0045] For the n One classifier, This represents a defined score threshold, and =( n -1) / , n =1, 2, ..., ,by p n Indicates the first n The prediction probabilities of each binary classifier, and the final quality score predicted by this branch is the prediction probabilities of each binary classifier. p ( x ref , x dist )=[ p 1, p 2, ..., p N The average value of ]. The MBC module will input pairs ( x ref , x dist Reference quality score Based on whether it is greater than the score threshold of each binary classifier Division, transformed A binary classification task was performed, and a new form was obtained. l ( x ref , x dist )=[ l 1, l 2, ..., l N ], for example, setting =5, and reference quality score =0.4, then the obtained quality fraction will be converted to y=[1, 1, 0, 0, 0]. In this way, the traditional quality score prediction task is transformed into a classification task, satisfying the above-described format of the generated data. The loss function chosen is BCELoss:
[0046] in P ( x ref , x dist ) represents the prediction vector for that branch. l ( x ref , x dist () represents the binary vector after label conversion. In this embodiment of the application, =5.
[0047] In one implementation, the image quality assessment network includes an anomaly type classification branch network. The step of inputting the target image and the distorted image into the image quality assessment network to obtain quality assessment scores and distortion types for the target image and the distorted image includes: inputting the target image and the distorted image into the image quality assessment network; extracting a first feature of the target image and a second feature of the distorted image; calculating the feature difference between the first feature and the second feature; and extracting and classifying the feature difference using a convolutional neural network and a multilayer perceptron of the anomaly type classification branch network to obtain the distortion type.
[0048] In this embodiment, the image quality assessment network further includes an anomaly type classification branch network, which inputs the target image and the distorted image into the image quality assessment network to extract the first feature Z of the target image. ref The second feature Z of the distorted image dist For the anomaly classification branch network, the feature difference |Z ref -Z dist After further extracting classification features through a convolutional neural network (CNN), the special vehicle is then extracted and classified through a multi-layer perceptron (MLP) module, thereby learning the distortion type of the target image and the distorted image. The loss function chosen is BCELoss.
[0049] For the anomaly type classification branch network, the special difference |Z ref -Z dist | This input pair is implemented by extracting and classifying classification features using a 5-layer CNN and MLP. x ref , x distThis task involves classifying distortion types (such as color cast, blur, occlusion, and black and white). The loss function uses BCELoss.
[0050] In this embodiment, the loss functions of the quality assessment branch network and the anomaly type classification branch network are weighted and summed to obtain the final loss function of the quality assessment network: = +
[0051] This application uses the aforementioned target image and the distorted image obtained by distorting the target image as part of the dataset. Based on the characteristics of the generation algorithm, the generated distorted image contains two pieces of information: distortion type and distortion degree. A higher distortion degree indicates a more distorted image, and the lower the quality score of this distorted image; the two are negatively correlated. Therefore, (1 - distortion level / total number of distortion levels) is used as the coarse label for the quality score in the quality evaluation branch. Additionally, the distortion type is used as the label for the anomaly type classification branch.
[0052] The specific implementation process of this application embodiment transforms the regression problem into a classification problem, thereby effectively utilizing the previously generated data and eliminating the need for additional manually labeled image quality data.
[0053] The network in this embodiment can perform quality assessment and anomaly classification of target and distorted images without requiring manual annotation of data quality scores. The quality assessment network first extracts the first feature of the target image and the second feature of the distorted image using a Siamese network, calculates the feature difference between them to obtain feature similarity, and then passes these features through a quality assessment branch network and an anomaly classification branch network. Based on fully utilizing the target image and the distorted image obtained after distortion transformation of the target image, it uses the distortion level as a coarse label to assess image quality, and simultaneously uses the distortion type to determine the anomaly category. The network design draws on the ideas of full-reference image quality assessment methods but overcomes the difficulties of full-reference methods with massive amounts of labeled data, greatly alleviating the pressure of data annotation.
[0054] The IQA algorithm proposed in this application uses the number of iterations and the type of distortion transformation as corresponding labels, eliminating the need for manual image annotation, thus saving significant manpower and greatly improving efficiency. It can evaluate the quality of a single image and determine its distortion type when a target image is available. The algorithm incorporates a Siamese network and uses a Conformer as the Encoder-Decoder module to further extract macroscopic and detailed image features, generating a quality evaluation branch. To intuitively display the distortion transformation type, a classification branch is also added to the network. The quality evaluation network ultimately outputs the quality score and distortion transformation type of the target object and the distorted image. It not only evaluates the image quality but also identifies the type of image anomaly, making the output richer.
[0055] The image quality assessment method provided in this application combines a Siamese network and uses a Conformer as the Encoder-Decoder module to further extract macroscopic and detailed features of the image, generating a quality assessment branch. To intuitively display the distortion transformation type, a classification branch is also added to the network. The network ultimately generates a quality score and distortion transformation type for each input pair. It not only evaluates the image quality but also identifies the type of image anomaly, enriching the output. It satisfies the requirement of manual annotation and labeled data with quality scores, enabling image quality assessment and distortion transformation type determination, greatly reducing the difficulty of manual data selection and improving efficiency.
[0056] It should be noted that the image quality evaluation method provided in this application embodiment can be executed by an image quality evaluation device or a control module within that device for executing the image quality evaluation method. This application embodiment uses an image quality evaluation device executing the image quality evaluation method as an example to illustrate the image quality evaluation device provided in this application embodiment.
[0057] Figure 6 This is a schematic diagram of the structure of an image quality evaluation device according to an embodiment of this application. Figure 6 As shown, the image quality evaluation device 600 includes a transformation module 610 and an evaluation module 620.
[0058] The transformation module 610 is used to perform a distortion transformation on the target image until the distorted image obtained by the distortion transformation of the target image meets the preset conditions; the evaluation module 620 is used to input the target image and the distorted image into an image quality evaluation network to obtain the quality evaluation score and distortion type of the target image and the distorted image.
[0059] In one implementation, the transformation module 610 is used to perform a distortion transformation on the target image using preset control parameters; if the similarity value between the obtained distorted image and the target image does not meet the preset conditions, the control parameters are adjusted, and the distorted image is distorted according to the adjusted control parameters until the similarity value between the distorted image and the target image meets the preset conditions, wherein the preset conditions include the similarity value between the distorted image and the target image being within a preset similarity range.
[0060] In one implementation, the image quality evaluation network includes a quality evaluation branch network. The evaluation module 620 is used to input the target image and the distorted image into the image quality evaluation network, extract a first feature of the target image and a second feature of the distorted image; calculate the feature difference between the first feature and the second feature; input the feature difference into the quality evaluation branch network, and output the quality evaluation scores of the target image and the distorted image.
[0061] In one implementation, the evaluation module 620 is configured to transform the feature difference through the decoder-encoder module of the quality evaluation branch network to obtain a difference image; input the difference image into a multi-binary classifier to determine a reference quality score for the difference image; determine the prediction probability that the score threshold of each binary classifier is less than or equal to the reference quality score based on the reference quality score and the score threshold of each binary classifier in the multi-binary classifier; and determine the average value of the prediction probabilities of each binary classifier as the quality evaluation score.
[0062] In one implementation, the evaluation module 620 is further configured to determine a score threshold for each binary classifier based on the number of each binary classifier in the multiple binary classifier.
[0063] In one implementation, the image quality evaluation network includes an anomaly type classification branch network. The evaluation module 620 is used to input the target image and the distorted image into the image quality evaluation network, extract a first feature of the target image and a second feature of the distorted image; calculate the feature difference between the first feature and the second feature; and extract and classify the feature difference through the convolutional neural network and multilayer perceptron of the anomaly type classification branch network to obtain the distortion type.
[0064] In one implementation, the distortion transformation includes at least one of the following: brightness transformation, color cast transformation, blur transformation, occlusion transformation, black and white transformation, screen tearing transformation, and noise transformation.
[0065] The image quality evaluation device in this application embodiment can be a device, or a component, integrated circuit, or chip in a terminal. The device can be a mobile electronic device or a non-mobile electronic device. For example, mobile electronic devices can be mobile phones, tablets, laptops, PDAs, in-vehicle electronic devices, wearable devices, ultra-mobile personal computers (UMPCs), netbooks, or personal digital assistants (PDAs), etc., while non-mobile electronic devices can be servers, network-attached storage (NAS), personal computers (PCs), televisions (TVs), ATMs, or self-service machines, etc. This application embodiment does not impose specific limitations.
[0066] The image quality evaluation device in this application embodiment can be a device with an operating system. This operating system can be Android, iOS, or other possible operating systems; this application embodiment does not specifically limit it.
[0067] The image quality evaluation device provided in this application embodiment can achieve... Figures 1 to 5 The various processes implemented in the method implementation examples will not be described again here to avoid repetition.
[0068] like Figure 7 As shown in the figure, this application embodiment also provides an electronic device 700, including a processor 701 and a memory 702. The memory 702 stores a program or instructions that can run on the processor 701. When the program or instructions are executed by the processor 701, they perform the following: distorting transformation on a target image until the distorted image obtained by distorting the target image meets preset conditions; inputting the target image and the distorted image into an image quality evaluation network to obtain the quality evaluation scores and distortion types of the target image and the distorted image.
[0069] In one implementation, the target image is distorted using preset control parameters. If the similarity value between the distorted image and the target image does not meet the preset conditions, the control parameters are adjusted, and the distorted image is distorted according to the adjusted control parameters until the similarity value between the distorted image and the target image meets the preset conditions, wherein the preset conditions include the similarity value between the distorted image and the target image being within a preset similarity range.
[0070] In one implementation, the image quality assessment network includes a quality assessment branch network. The target image and the distorted image are input into the image quality assessment network to extract a first feature of the target image and a second feature of the distorted image. The feature difference between the first feature and the second feature is calculated. The feature difference is input into the quality assessment branch network to output the quality assessment scores of the target image and the distorted image.
[0071] In one implementation, the feature difference is transformed by the decoder-encoder module of the quality evaluation branch network to obtain a difference image; the difference image is input into a multi-binary classifier to determine a reference quality score; based on the reference quality score and the score threshold of each binary classifier of the multi-binary classifier, a prediction probability is determined that the score threshold of each binary classifier is less than or equal to the reference quality score; the average value of the prediction probabilities of each binary classifier is determined as the quality evaluation score.
[0072] In one implementation, after inputting the difference image into a multi-binary classifier and determining a reference quality score for the difference image, a score threshold for each binary classifier is determined based on the number of each binary classifier in the multi-binary classifier.
[0073] In one implementation, the image quality assessment network includes an anomaly type classification branch network. The target image and the distorted image are input into the image quality assessment network to extract a first feature of the target image and a second feature of the distorted image. The feature difference between the first feature and the second feature is calculated. The feature difference is extracted and classified by the convolutional neural network and multilayer perceptron of the anomaly type classification branch network to obtain the distortion type.
[0074] In one implementation, the distortion transformation includes at least one of the following: brightness transformation, color cast transformation, blur transformation, occlusion transformation, black and white transformation, screen tearing transformation, and noise transformation.
[0075] The specific execution steps can be found in the various steps of the above-described image quality evaluation method embodiments, and can achieve the same technical effect. To avoid repetition, they will not be repeated here.
[0076] It should be noted that the electronic devices in the embodiments of this application include: servers, terminals, or other devices besides terminals.
[0077] The above electronic device structure does not constitute a limitation on the electronic device. An electronic device may include more or fewer components than illustrated, or combine certain components, or arrange them differently. For example, an input unit may include a Graphics Processing Unit (GPU) and a microphone, and a display unit may use a liquid crystal display (LCD), organic light-emitting diode (OLED), or other similar display panels. User input units include at least one of a touch panel and other input devices. A touch panel is also called a touchscreen. Other input devices may include, but are not limited to, physical keyboards, function keys (such as volume control buttons, power buttons, etc.), trackballs, mice, and joysticks, which will not be elaborated further here.
[0078] Memory can be used to store software programs and various data. Memory can primarily include a first storage area for storing programs or instructions and a second storage area for storing data. The first storage area can store the operating system, application programs or instructions required for at least one function (such as sound playback, image playback, etc.). Furthermore, memory can include volatile memory or non-volatile memory, or both. Non-volatile memory can be read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), or flash memory. Volatile memory can be random access memory (RAM), static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDRSDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous linked dynamic random access memory (Synchlink DRAM, SLDRAM), and direct memory bus RAM (DRRAM).
[0079] The processor may include one or more processing units; optionally, the processor integrates an application processor and a modem processor, wherein the application processor mainly handles operations related to the operating system, user interface, and applications, while the modem processor mainly handles wireless communication signals, such as a baseband processor. It is understood that the aforementioned modem processor may also not be integrated into the processor.
[0080] This application also provides a readable storage medium storing a program or instructions. When the program or instructions are executed by a processor, they implement the various processes of the above-described image quality evaluation method embodiments and achieve the same technical effect. To avoid repetition, they will not be described again here.
[0081] The processor is the processor in the electronic device described in the above embodiments. The readable storage medium includes computer-readable storage media, such as ROM, RAM, magnetic disk, or optical disk.
[0082] This application also provides a computer program product, which includes a computer program stored on a non-transitory computer-readable storage medium. The computer program includes program instructions, which, when executed by a computer, cause the computer to perform various processes of the above-described image quality evaluation method embodiments and achieve the same technical effect. To avoid repetition, these will not be described again here.
[0083] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element. Furthermore, it should be noted that the scope of the methods and apparatuses in the embodiments of this application is not limited to performing functions in the order shown or discussed, but may also include performing functions substantially simultaneously or in the reverse order, depending on the functions involved. For example, the described methods may be performed in a different order than described, and various steps may be added, omitted, or combined. Additionally, features described with reference to certain examples may be combined in other examples.
[0084] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a computer software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions to cause a terminal (which may be a mobile phone, computer, server, or network device, etc.) to execute the methods described in the various embodiments of this application.
[0085] The embodiments of this application have been described above with reference to the accompanying drawings. However, this application is not limited to the specific embodiments described above. The specific embodiments described above are merely illustrative and not restrictive. Those skilled in the art can make many other forms under the guidance of this application without departing from the spirit and scope of the claims, and all of these forms are within the protection scope of this application.
Claims
1. An image quality assessment method, characterized in that, include: The target image is distorted until the distorted image obtained by distorting the target image meets the preset conditions; The target image and the distorted image are input into an image quality assessment network to obtain the quality assessment scores and distortion types of the target image and the distorted image; The image quality assessment network includes a quality assessment branch network. The process of inputting the target image and the distorted image into the image quality assessment network to obtain the quality assessment scores and distortion types of the target image and the distorted image includes: The target image and the distorted image are input into an image quality assessment network to extract the first feature of the target image and the second feature of the distorted image. Calculate the feature difference between the first feature and the second feature; The feature difference is input into the quality evaluation branch network, and the quality evaluation scores of the target image and the distorted image are output. The step of inputting the feature difference into the quality evaluation branch network and outputting the quality evaluation scores of the target image and the distorted image includes: The feature difference is transformed by the decoder-encoder module of the quality evaluation branch network to obtain a difference image; The difference image is input into a multi-binary classifier to determine a reference quality score for the difference image; Based on the reference quality score and the score threshold of each binary classifier of the multi-binary classifier, determine the predicted probability that the score threshold of each binary classifier is less than or equal to the reference quality score; The average predicted probability of each of the binary classifiers is determined as the quality evaluation score.
2. The method according to claim 1, characterized in that, The process of distorting the target image until the distorted image obtained by the distortion transformation meets preset conditions includes: The target image is distorted using preset control parameters; If the similarity value between the obtained distorted image and the target image does not meet the preset conditions, the control parameters are adjusted, and the distorted image is distorted according to the adjusted control parameters until the similarity value between the distorted image and the target image meets the preset conditions. The preset conditions include that the similarity value between the distorted image and the target image is within a preset similarity range.
3. The method according to claim 1, characterized in that, After inputting the difference image into a multi-binary classifier to determine the reference quality score of the difference image, the method further includes: The score threshold for each binary classifier is determined based on the number of each binary classifier in the multiple binary classifier.
4. The method according to claim 1, characterized in that, The image quality assessment network includes an anomaly type classification branch network. The process of inputting the target image and the distorted image into the image quality assessment network to obtain the quality assessment scores and distortion types of the target image and the distorted image includes: The target image and the distorted image are input into an image quality assessment network to extract the first feature of the target image and the second feature of the distorted image. Calculate the feature difference between the first feature and the second feature; The feature difference is extracted and classified by the convolutional neural network and multilayer perceptron of the anomaly type classification branch network to obtain the distortion type.
5. The method according to claim 1, characterized in that, The distortion transformation includes at least one of the following: brightness transformation, color cast transformation, blur transformation, occlusion transformation, black and white transformation, screen distortion transformation, and noise transformation.
6. An electronic device, characterized in that, It includes a processor, a memory, and a program or instructions stored in the memory and executable on the processor, wherein the program or instructions, when executed by the processor, implement the steps of the image quality evaluation method as described in any one of claims 1-5.
7. A readable storage medium, characterized in that, The readable storage medium stores a program or instructions that, when executed by a processor, implement the steps of the image quality evaluation method as described in any one of claims 1-5.
8. A computer program product, characterized in that, The computer program product includes a computer program stored on a non-transitory computer-readable storage medium, the computer program including program instructions that, when executed by a computer, cause the computer to perform the steps of the image quality evaluation method as described in any one of claims 1-5.
Citation Information
Patent Citations
Objective evaluation method for full reference image quality based on neural network learning integration
CN108615231A
Face image quality evaluation model construction method and device, equipment and medium
CN113505854A