Image quality evaluation method and system
By comparing the network and deep network jointly extracting features and using adaptive weight attention mechanism fusion, the problem of insufficient image quality evaluation in the prior art is solved, and the problem of inadequate global context modeling capabilities and inconvolution kernel weight allocation is achieved, achieving higher evaluation accuracy and adaptability.
Patent Information
- Application Number
- CN202510607827.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-13
- Publication Date
- 2025-06-10
AI Technical Summary
In the image quality evaluation, the prior art has problems such as insufficient modeling ability of the global context of the image and the difficulty in adapting to dynamic combination of multiple types of distortions.
An image quality evaluation method is proposed. By extracting the first feature of robust distinction by comparing the network, and combining the second feature of image texture, edge, shape and other information in the deep network, the adaptive weight attention mechanism is used to fuse the two to evaluate the image quality.
Improves adaptability and evaluation accuracy to distortion types, allowing more efficient handling of uncertainty in mixed distortion types and degrees.
Smart Images

Figure CN120125975A_ABST
Abstract
Description
Technical Field
[0001] The embodiments of the present application relate to the technical field of distorted image processing, and in particular, to an image quality assessment method and system. Background Art
[0002] Before the human eye receives visual information of an image, multimedia data usually needs to go through the entire information processing chain, and various perceptual quality degradations will occur at different processing stages. During the acquisition process, defocusing and blurring may be caused by device jitter, block effects and ringing effects may be generated during compression, time delay, error codes, and data packet loss may occur during transmission, and there may be noise during the reconstruction process, resulting in the image finally displayed to the viewer being degraded. Image quality is easily affected by multiple factors such as noise, blurring, and compression distortion in imaging devices, transmission processes, and post-processing. Its accurate assessment is crucial for computer vision applications.
[0003] Image Quality Assessment (IQA) is a fundamental issue in the fields of image processing, image or video coding, etc., mainly used to evaluate the distortion degree of images or videos, and plays an important role in computer vision, image processing, and multimedia applications.
[0004] In recent years, deep learning technologies have promoted the innovation of image quality assessment methods. Models represented by convolutional neural networks (CNNs) and Transformers have significantly improved the quality prediction accuracy through end-to-end feature learning. CNNs rely on convolutional kernel weight sharing and local receptive field design, use convolutional layers to extract spatial texture features, pooling layers to compress feature dimensions, and fully connected layers to regress quality scores. Its built-in inductive bias characteristics greatly reduce the number of model parameters. However, the strong local preference of CNNs leads to insufficient ability to model the global context of images, and the fixed convolutional kernel weight allocation is difficult to adapt to the dynamic combination of multiple types of distortions. Although CNN-based methods attempt to improve the evaluation effect through distortion classification, in the face of the uncertainty of mixed distortion types and degrees in real scenarios, there is still a lack of effective adaptive constraint mechanisms. Summary of the Invention
[0005] The following is an overview of the subject matter described in detail in this document. This overview is not intended to limit the scope of protection of the claims.
[0006] The main purpose of the embodiments of the present disclosure is to propose an image quality assessment method and system that can improve the accuracy of quality assessment of distorted images.
[0007] The first aspect of the embodiments of the present application proposes an image quality assessment method, and the method includes: Obtain an image to be evaluated; Extract the first feature of the image to be evaluated according to the contrast network; wherein, the contrast network is obtained by contrast training between positive sample pairs and negative sample pairs based on training data, the training data includes basic distorted images and positive and negative samples transformed from the basic distorted images, the distortion type between the positive sample and the basic distorted image is the same, the distortion type between the negative sample and the basic distorted image is different, the positive sample pair is composed of the basic distorted image and the positive sample, and the negative sample pair is composed of the basic distorted image and the negative sample; Extract the second feature of the image to be evaluated according to the deep network; Fuse the first feature and the second feature of the image to be evaluated through an adaptive weight attention mechanism to obtain a fused feature; Evaluate the quality of the image to be evaluated based on the fused feature.
[0008] An image quality evaluation method provided by an embodiment of the present disclosure has at least the following beneficial effects: On the one hand, this method obtains a contrast network through contrast training between positive sample pairs and negative sample pairs based on training data. Training the contrast network in this way can make similar images with the same distortion type closer in the feature space, and the distinguishability of images with different distortion types in the feature space is greater. Furthermore, the contrast network can extract more robust and distinguishable first features in the image to be evaluated. On the other hand, directly use the deep network to extract the second features with information such as image texture, edges, and shapes in the image to be evaluated. Finally, fuse the first feature and the second feature based on the attention mechanism. Finally, evaluate the quality of the image to be evaluated through the fused feature. Here, the first feature and the second feature are adaptively fused through the attention mechanism, which can adaptively give more attention to the region features in the image that are most significantly affected by distortion, thereby improving the adaptability to the distortion type and ultimately enhancing the accuracy of evaluating the quality of the image to be evaluated.
[0009] In some embodiments, the contrast network includes an encoder and a projection head; The training process of the contrast network includes: Extract the respective deep features of the basic distorted image, the positive sample, and the negative sample through the encoder, and extract the respective projection features from the respective deep features according to the projection head; Calculate the first similarity of the projection features between the positive sample pairs, and calculate the second similarity of the projection features between the negative sample pairs; Calculate a contrast loss based on the first similarity and the second similarity; Perform backpropagation on the encoder and the projection head based on the contrast loss until the training is completed.
[0010] In some embodiments, calculating a contrast loss based on the first similarity and the second similarity includes: ; Wherein, is the contrast loss, is the natural exponential function, is the projection feature of the base distortion image in the positive sample pair and the projection feature of the positive sample The first similarity between is the projection feature of the base distortion image in the negative sample pair and the projection feature of the negative sample The second similarity between is the temperature parameter.
[0011] In some embodiments, the deep network includes a global feature extraction network and a local feature extraction network; Extracting the second feature of the image to be evaluated according to the deep network includes: Extracting the global feature of the image to be evaluated according to the global feature extraction network; Extracting the local feature of the image to be evaluated according to the local feature extraction network; Fusing the global feature and the local feature to obtain the second feature of the image to be evaluated.
[0012] In some embodiments, the global feature extraction network is a Transformer network, and the local feature extraction network is a CNN network.
[0013] In some embodiments, evaluating the quality of the image to be evaluated based on the fusion feature includes: Regressing the fusion feature through a global average pooling layer to obtain the quality of the image to be evaluated; The loss function during the training process of the contrast network, the deep network, and the global average pooling layer is: ; Wherein, is the number of base distortion images, is the quality of the th base distortion image output by the global average pooling layer, is the th actual quality of the base distortion image.
[0014] In some embodiments, fusing the first feature and the second feature of the image to be evaluated by an adaptive weight attention mechanism to obtain a fusion feature includes: Map the first feature of the image to be evaluated into a query vector; Map the second feature of the image to be evaluated into key and value vectors; According to the key and value vectors and the query vector, calculate the fused feature by using the adaptive weighted attention mechanism formula.
[0015] In a second aspect of the embodiments of the present application, an image quality evaluation system is proposed. The system includes: An image acquisition module, configured to acquire an image to be evaluated; A first calculation module, configured to extract the first feature of the image to be evaluated according to a contrast network; wherein, the contrast network is obtained by performing contrast training between positive sample pairs and negative sample pairs based on training data, the training data includes a basic distorted image and positive and negative samples obtained by transforming the basic distorted image, the distortion type between the positive sample and the basic distorted image is the same, the distortion type between the negative sample and the basic distorted image is different, the positive sample pair is composed of the basic distorted image and the positive sample, and the negative sample pair is composed of the basic distorted image and the negative sample; A second calculation module, configured to extract the second feature of the image to be evaluated according to a deep network; A feature fusion module, configured to perform adaptive weight attention mechanism fusion on the first feature and the second feature of the image to be evaluated to obtain a fused feature; A quality evaluation module, configured to evaluate the quality of the image to be evaluated based on the fused feature.
[0016] In a third aspect of the embodiments of the present application, an electronic device is proposed, including at least one controller and a memory communicatively connected to the at least one controller; the memory stores instructions executable by the at least one controller, and when the instructions are executed by the controller, the controller is caused to execute the image quality evaluation method as described in the first aspect.
[0017] In a fourth aspect of the embodiments of the present application, a computer-readable storage medium is proposed. The computer-readable storage medium stores computer-executable instructions, and the computer-executable instructions are used to cause a computer to execute the image quality evaluation method as described above.
[0018] It can be understood that the beneficial effects of the above second aspect to the fourth aspect compared with the related art are the same as those of the above first aspect compared with the related art. For the relevant descriptions, reference can be made to the relevant descriptions in the above first aspect, and details are not described herein again. Description of the Drawings
[0019] To more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the accompanying drawings required for use in the embodiments or the description of related technologies. Obviously, the accompanying drawings in the following description are only some embodiments of the embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other accompanying drawings can also be obtained based on these drawings.
[0020] Figure 1 is a flowchart of an embodiment of an image quality assessment method provided by the present application; Figure 2 is a schematic structural diagram of an image quality assessment model provided by the present application; Figure 3 is a structural diagram of an embodiment of an image quality assessment system provided by the present application; Figure 4 is a structural diagram of an embodiment of an electronic device provided by the present application. Detailed implementation manners
[0021] In order to make the purpose, technical solutions and advantages of the present application clearer, the following further details the present application in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.
[0022] It should be noted that although functional module division is carried out in the device schematic diagram and the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order from the module division in the device or the flowchart. Terms such as "first" and "second" in the specification, claims and the above accompanying drawings are used to distinguish similar objects and do not necessarily need to describe a specific order or sequence.
[0023] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which this application belongs. The terms used herein are only for the purpose of describing the embodiments of this application and are not intended to limit this application. Before the human eye receives the visual information of the image, the multimedia data usually needs to go through the entire information processing chain, and various perceptual quality degradations will occur at different processing stages. During the acquisition process, defocus and blur may be caused by device jitter, block effects and ringing effects may occur during compression, time delay, error codes and data packet loss may occur during transmission, and noise may occur during the reconstruction process, resulting in the image finally displayed to the viewer being degraded. Image quality is vulnerable to multiple factors such as noise, blur, and compression distortion in the imaging device, transmission process and post-processing, and its accurate assessment is crucial for computer vision applications.
[0024] Image Quality Assessment (IQA) is a fundamental issue in the fields of image processing, image or video coding, etc. It is mainly used to evaluate the distortion degree of images or videos and plays an important role in computer vision, image processing, and multimedia applications.
[0025] In recent years, deep learning technologies have promoted the innovation of image quality assessment methods. Models represented by convolutional neural networks (CNNs) and Transformers have significantly improved the quality prediction accuracy through end-to-end feature learning. With the design of convolutional kernel weight sharing and local receptive fields, CNNs use convolutional layers to extract spatial texture features, pooling layers to compress the feature dimensions, and fully connected layers to regress the quality scores. Its built-in inductive bias characteristics greatly reduce the number of model parameters. However, the strong local preference of CNNs leads to insufficient ability to model the global context of images, and the fixed convolutional kernel weight allocation is difficult to adapt to the dynamic combination of multiple types of distortions. Although CNN-based methods attempt to improve the evaluation effect through distortion classification, in the face of the uncertainty of mixed distortion types and degrees in real scenarios, there is still a lack of effective adaptive constraint mechanisms.
[0026] To solve the above technical problems, such as Figure 1 , an embodiment of this application provides an image quality assessment method, which includes steps S110 to S150: Step S110, obtain the image to be evaluated.
[0027] In step S110, the image to be evaluated refers to the relevant image that needs to be quality-evaluated, and there is no limit to the specific type. For example, after an original image goes through the entire information processing chain, during the acquisition process of the original image, it may be out of focus or blurred due to device jitter, block effects or ringing effects may occur during compression, time delays, bit errors, data packet losses may occur during transmission, and there may be noise during the reconstruction process, resulting in the image to be evaluated finally displayed to the viewer being a distorted image. Therefore, it is necessary to perform quality assessment on the image to be evaluated.
[0028] Step S120, extract the first feature of the image to be evaluated according to the contrast network; among them, the contrast network is obtained through contrast training between positive sample pairs and negative sample pairs based on training data.
[0029] In step S120, the training data includes basic distorted images and positive and negative samples obtained by transforming the basic distorted images (for example, adding noise). The distortion type between the positive samples and the basic distorted images is the same, and the distortion type between the negative samples and the basic distorted images is different. The positive sample pair consists of the basic distorted image and the positive sample, and the negative sample pair consists of the basic distorted image and the negative sample.
[0030] The image quality assessment model of this application mainly includes two networks, namely, a contrast network and a deep network. Here, the contrast network is mainly introduced: The contrast network can be composed of an encoder (such as ResNet-50), or composed of ResNet-50 and a projection head. The purpose of adding the projection head here is to extract projection features for subsequent calculations. Taking the latter as an example, based on ResNet-50, the deep learning features of the base distorted image, positive samples, and negative samples can be extracted respectively. Then, the projection head can extract the respective projection features from the respective deep learning features. Furthermore, the contrast learning mechanism can be used to optimize the feature space, making similar images (i.e., positive sample pairs) closer in the feature space and dissimilar images (i.e., negative sample pairs) farther away in the feature space. Thus, the contrast network can learn discriminative features. Then, when the image to be evaluated is input into the trained contrast network, more robust and discriminative first features in the image to be evaluated can be learned.
[0031] Step S130: Extract the second feature of the image to be evaluated according to the deep network.
[0032] In step S130, the deep network can be a CNN, a Transformer, or a combination of a CNN and a Transformer. The deep network can directly extract the second feature with information such as image texture, edges, and shapes in the image to be evaluated.
[0033] Step S140: Fuse the first feature and the second feature of the image to be evaluated through an adaptive weight attention mechanism to obtain a fused feature.
[0034] In step S140, by adaptively fusing the first feature and the second feature through the attention mechanism, more attention can be adaptively given to the region features in the image that are most significantly affected by distortion.
[0035] Step S150: Evaluate the quality of the image to be evaluated based on the fused feature.
[0036] In step S150, a score corresponding to the fused feature can be given through, for example, a global average pooling layer.
[0037] The method provided by this application has at least the following beneficial effects: On the one hand, this method performs contrast training between positive sample pairs and negative sample pairs based on training data to obtain a contrast network. Training the contrast network in this way can make similar images of the same distortion type closer in the feature space, and the images of different distortion types have greater distinguishability in the feature space. Furthermore, the contrast network can extract more robust and discriminative first features in the image to be evaluated. On the other hand, a deep network is directly used to extract second features with information such as image texture, edges, and shapes in the image to be evaluated. Finally, the first feature and the second feature are fused based on an attention mechanism, and the quality of the image to be evaluated is evaluated through the fused features. Here, the first feature and the second feature are adaptively fused through the attention mechanism, which can adaptively pay more attention to the features of the regions in the image that are most significantly affected by distortion, thereby improving the adaptability to distortion types and ultimately enhancing the accuracy of evaluating the quality of the image to be evaluated.
[0038] Furthermore, the contrast network includes an encoder and a projection head; The training process of the contrast network includes steps S210 to S240: In step S210, the encoder extracts the respective deep features of the base distortion image, positive samples, and negative samples, and the projection features are extracted from the respective deep features according to the projection head.
[0039] In step S210, the encoder can be a ResNet-50 network.
[0040] In step S220, the first similarity of the projection features between positive sample pairs is calculated, and the second similarity of the projection features between negative sample pairs is calculated.
[0041] In step S220, calculations can be performed using, for example, the Euclidean distance.
[0042] In step S230, the contrast loss is calculated based on the first similarity and the second similarity.
[0043] In step S230, the contrast loss can be calculated through the following formula: ; where is the contrast loss, is the natural exponential function, is the projection feature of the base distortion image in the positive sample pair and the projection feature of the positive sample the first similarity between them, is the projection feature of the base distortion image in the negative sample pair and the projection feature of the negative sample the second similarity between them, is the temperature parameter.
[0044] Step S240: Perform backpropagation on the encoder and the projection head based on the contrastive loss until the training is completed.
[0045] Finally, train the model through backpropagation. When the number of training times is reached, the training is completed.
[0046] In this embodiment, the contrastive network is obtained by performing contrastive training between positive sample pairs and negative sample pairs based on training data. Training the contrastive network in this way can make similar images of the same distortion type closer in the feature space, and the distinguishability of images of different distortion types in the feature space is greater. Furthermore, the contrastive network can extract more robust discriminative features in the images.
[0047] Furthermore, the deep network includes a global feature extraction network and a local feature extraction network.
[0048] Extracting the second feature of the image to be evaluated according to the deep network in step S130 includes the following steps S310 to S330: Step S310: Extract the global feature of the image to be evaluated according to the global feature extraction network.
[0049] In this step, the global feature reflects the overall attributes or statistical information in the image, such as color histogram, texture distribution, etc.
[0050] Step S320: Extract the local feature of the image to be evaluated according to the local feature extraction network.
[0051] In this step, the local feature contains more detailed features extracted from specific regions of the image (such as key points, edges).
[0052] Step S330: Obtain the second feature of the image to be evaluated according to the fusion of the global feature and the local feature.
[0053] In step 130, by fusing the global feature and the local feature, more feature information can be learned to improve the accuracy of image evaluation.
[0054] Furthermore, the global feature extraction network is a Transformer network, and the local feature extraction network is a CNN network. With the design of convolutional kernel weight sharing and local receptive fields, the CNN can extract spatial texture features in the convolutional layer and compress the feature dimensions in the pooling layer, and can extract more detailed features. The Transformer network can extract the global feature of the image.
[0055] Furthermore, evaluating the quality of the image to be evaluated based on the fusion feature in step S150 includes the following step S510: Step S510: Perform regression on the fused features through a global average pooling layer to obtain the quality of the image to be evaluated. The loss functions during the training processes of the comparison network, the deep network, and the global average pooling layer are as follows: ; Among them, is the number of base distorted images, is the quality of the th base distorted image output by the global average pooling layer, is the th actual quality of the base distorted image.
[0056] This method uses minimizing the regression loss to train the quality assessment model. Here, by minimizing the regression loss, the evaluation accuracy of the model can be improved.
[0057] Furthermore, fusing the first feature and the second feature of the image to be evaluated in step S140 through an adaptive weight attention mechanism to obtain fused features includes the following steps S410 to S430: Step S410: Map the first feature of the image to be evaluated into a query vector.
[0058] Step S420: Map the second feature of the image to be evaluated into key and value vectors.
[0059] Step S430: Calculate the fused features according to the key and value vectors and the query vector using the adaptive weighted attention mechanism formula.
[0060] The calculation formula of this method includes: ; Among them, is the query vector, is the key vector, is the value vector, is the scaling factor to prevent gradient vanishing, is the transpose.
[0061] As Figure 2 , for ease of understanding, an image quality assessment method is provided. This method includes the following steps: Step S910: Construct an image quality assessment model and train it.
[0062] As Figure 2The model shown includes: a deep network (part for extracting feature R), a contrast network (part for extracting feature P), a patch attention block, and a global average pooling; the deep network includes a parallel CNN and Transformer. The contrast network includes an encoder (such as ResNet-50) and a projection head (such as a multi-layer perceptron).
[0063] Select a high-quality base distortion image , and generate positive samples based on the base distorted image , generate negative samples based on the base distorted image Then construct the positive sample pair And negative sample pairs Among them, the positive sample With the base distorted image The same distortion type, negative samples With the base distorted image The distortion type is different.
[0064] Then, the training images are extracted through the encoder , positive sample and negative samples The deep features , , .
[0065] Then, the depth features are extracted through the projection head , , Corresponding projection features , , .
[0066] Then, calculate the positive sample pair Cosine similarity of , and negative samples The cosine similarity of .
[0067] Then, the loss is calculated based on the formula: ; in, is the contrast loss, is a natural exponential function, and the temperature parameter .
[0068] Finally, back propagation is performed to optimize the feature space by contrast loss, so that similar images are closer in the feature space, thereby learning discriminative features.
[0069] Step S920: extracting the first feature of the image to be evaluated based on the contrast network 。 As one of the inputs of the patch attention block.
[0070] Step S930, extract local features of the image to be evaluated based on CNN 。
[0071] Step S940, extract global features of the image to be evaluated based on Transformer 。
[0072] Step S950, fuse the local features and the global features to obtain the second feature Feature As another input.
[0073] Step S960, input the first feature into the patch attention block as a query vector. Additionally, use the second feature as key and value vectors, and calculate through the following formula: ; Obtain the fused feature 。
[0074] Step S970, regress the quality score using the quality vector obtained through global average pooling (GAP).
[0075] During the training process of the model, minimize the regression loss to obtain the optimal training result: ; where is the number of base distortion images, is the quality of the th base distortion image in the global average pooling output, is the th actual quality of the base distortion image.
[0076] The beneficial effect of this method is that: During the training process of the contrast network, the method first performs a series of random transformations on the base distortion map to generate positive sample pairs with the same distortion type and negative sample pairs with different distortion types, and uses an encoder (ResNet50) to extract features from the transformed images to obtain high-dimensional feature representations. Then, by calculating the cosine similarity between the feature vectors and optimizing the contrast loss, similar images are made closer in the feature space, thereby learning more robust discriminative features. In the usage stage of the model, the image to be evaluated first passes through a multi-level feature extraction module of CNN and transformer, gradually capturing key information such as the texture, edges, and shapes of the image, and performing local region sampling on feature maps at different levels to enhance the model's perception ability of local distortion. Subsequently, the first feature and the second feature are weighted by the patch attention block, which uses the attention mechanism to adaptively focus on the regions in the image that are most significantly affected by the distortion. Images with different distortion types have different responses in the feature space, thereby improving the adaptability to the distortion type. The features after attention weighting are further subjected to quality scoring, and the final result of the image quality assessment is calculated based on the learned feature representation and normalized to between [0,1] to ensure the interpretability and consistency of the scoring.
[0077] This method introduces contrast learning and the attention mechanism into image quality assessment. First, positive and negative sample pairs are constructed based on contrast learning to optimize the feature space, making the distinction between high-quality and low-quality images greater in the feature space and improving robustness. The attention mechanism enhances the feature representation of the significantly distorted regions by adaptively allocating weights, combining the low-level edge features extracted by CNN and the high-level content features of the image extracted by Transformer, making the evaluation more accurate. Combining the two can make full use of the structural information and local details of the image and improve the adaptive ability of the model.
[0078] As Figure 3 shown, an image quality assessment system is provided, and the system includes: An image acquisition module 1100 is used to acquire the image to be evaluated; A first calculation module 1200 is used to extract the first feature of the image to be evaluated according to the contrast network. Among them, the contrast network is obtained by performing contrast training between positive and negative sample pairs based on the training data. The training data includes the base distortion images and the positive and negative samples obtained by transforming the base distortion images. The distortion type between the positive sample and the base distortion image is the same, and the distortion type between the negative sample and the base distortion image is different. The positive sample pair is composed of the base distortion image and the positive sample, and the negative sample pair is composed of the base distortion image and the negative sample; The second computing module 1300 is used to extract the second feature of the image to be evaluated according to the deep network; The feature fusion module 1400 is used to perform adaptive weight attention mechanism fusion on the first feature and the second feature of the image to be evaluated to obtain a fused feature; The quality evaluation module 1500 is used to evaluate the quality of the image to be evaluated based on the fused feature.
[0079] It should be noted that the image quality evaluation system provided in this embodiment and the above image quality evaluation method are based on the same inventive concept. Therefore, the relevant content of the above image quality evaluation method also applies to the content of the image quality evaluation system. Therefore, it will not be elaborated here.
[0080] As Figure 4 shown, an embodiment of the present application also provides an electronic device, which includes: At least one memory; At least one processor; At least one program; The program is stored in the memory, and the processor executes at least one program to implement the above-mentioned image quality evaluation method of the present disclosure.
[0081] The electronic device can be any intelligent terminal including a mobile phone, a tablet computer, a personal digital assistant (Personal Digital Assistant, PDA), a vehicle-mounted computer, etc.
[0082] The electronic device of the embodiment of the present application will be introduced in detail below.
[0083] The processor 1600 can be implemented by using a general-purpose central processing unit (Central Processing Unit, CPU), a microprocessor, an application-specific integrated circuit (Application Specific Integrated Circuit, ASIC), or one or more integrated circuits, etc., and is used to execute relevant programs to implement the technical solution provided by the embodiment of the present application; The memory 1700 can be implemented in the form of a read-only memory (Read Only Memory, ROM), a static storage device, a dynamic storage device, or a random access memory (Random Access Memory, RAM). The memory 1700 can store an operating system and other application programs. When implementing the technical solution provided by the embodiments of this specification through software or firmware, the relevant program codes are stored in the memory 1700, and the processor 1600 is called to execute the image quality evaluation method of the embodiment of the present application.
[0084] An input / output interface 1800 for implementing information input and output; A communication interface 1900 for implementing communication interaction between this device and other devices, which can achieve communication through wired means (such as USB, network cable, etc.) or wireless means (such as mobile network, WIFI, Bluetooth, etc.); A bus 2000 for transmitting information between various components of the device (such as the processor 1600, the memory 1700, the input / output interface 1800, and the communication interface 1900); Among them, the processor 1600, the memory 1700, the input / output interface 1800, and the communication interface 1900 are communicatively connected to each other inside the device through the bus 2000.
[0085] The embodiment of the present application also provides a storage medium, which is a computer-readable storage medium, and the computer-readable storage medium stores computer-executable instructions for causing a computer to execute the above image quality assessment method.
[0086] As a non-transitory computer-readable storage medium, the memory can be used to store non-transitory software programs and non-transitory computer-executable programs. In addition, the memory may include high-speed random access memory, and may also include non-transitory memory, such as at least one magnetic disk storage device, a flash memory device, or other non-transitory solid-state storage devices.
[0087] In some embodiments, the memory may optionally include a memory remotely provided with respect to the processor, and these remote memories can be connected to the processor through a network. Examples of the above network include but are not limited to the Internet, an enterprise intranet, a local area network, a mobile communication network, and combinations thereof.
[0088] The embodiments described in the present application are for more clearly illustrating the technical solutions of the embodiments of the present application, and do not constitute a limitation on the technical solutions provided by the embodiments of the present application. Those skilled in the art can know that with the evolution of technology and the emergence of new application scenarios, the technical solutions provided by the embodiments of the present application are equally applicable to similar technical problems.
[0089] Those skilled in the art can understand that the technical solutions shown in the figures do not constitute a limitation on the embodiments of the present application, and may include more or fewer steps than those shown, or combine certain steps, or different steps.
[0090] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, that is, they may be located in one place, or may be distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0091] Those of ordinary skill in the art can understand that all or some of the steps in the methods disclosed above, and the functional modules / units in systems and devices, can be implemented as software, firmware, hardware, or a suitable combination thereof.
[0092] As used in the specification of this application and the above-mentioned drawings, terms such as "first", "second", "third", "fourth", etc. (if any) are used to distinguish similar objects and do not necessarily describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments of this application described here can be implemented in an order different from those illustrated or described here. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products, or devices.
[0093] It should be understood that in this application, "at least one (item)" means one or more, and "a plurality" means two or more. "And / or" is used to describe the association relationship of associated objects and indicates that there can be three relationships. For example, "A and / or B" can mean: only A exists, only B exists, and both A and B exist at the same time. Among them, A and B can be singular or plural. The character " / " generally means that the associated objects before and after are in an "or" relationship. "At least one (one) of the following" or a similar expression means any combination of these items, including any combination of single items (ones) or plural items (ones). For example, at least one (one) of a, b, or c can mean: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, and c can be single or multiple.
[0094] In several embodiments provided in this application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of units is only a logical function division, and there can be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling, direct coupling, or communication connection to each other can be through some interfaces, and the indirect coupling or communication connection of devices or units can be in electrical, mechanical, or other forms.
[0095] The unit described as a separation component may or may not be physically separated. The component displayed as a unit may or may not be a physical unit, that is, it may be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0096] In addition, each functional unit in various embodiments of the present application may be integrated in a processing unit, may exist separately as individual physical units, or two or more units may be integrated in one unit. The above integrated unit can be implemented in the form of hardware or in the form of a software functional unit.
[0097] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes multiple instructions for causing an electronic device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods in various embodiments of the present application. The aforementioned storage medium includes: various media that can store programs such as USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs.
[0098] The above is a specific description of the preferred implementation of the embodiments of the present application. However, the embodiments of the present application are not limited to the above implementation manners. Those skilled in the art can make various equivalent deformations or substitutions without departing from the spirit of the embodiments of the present application. These equivalent deformations or substitutions are all included within the scope defined by the claims of the embodiments of the present application.
Claims
1. A method for evaluating image quality, characterized in that: The method comprises: Obtaining an image to be evaluated; Extracting the first feature of the image to be evaluated according to a contrast network; wherein the contrast network is obtained by contrast training between positive sample pairs and negative sample pairs based on training data, the training data includes a basic distorted image and positive samples and negative samples obtained by transformation based on the basic distorted image, the positive sample and the basic distorted image have the same distortion type, the negative sample and the basic distorted image have different distortion types, the positive sample pair is composed of the basic distorted image and the positive sample, and the negative sample pair is composed of the basic distorted image and the negative sample; Extracting a second feature of the image to be evaluated according to a deep network; The first feature and the second feature of the image to be evaluated are fused by an adaptive weighted attention mechanism to obtain a fused feature; The quality of the image to be evaluated is evaluated based on the fusion feature.
2. The image quality assessment method according to claim 1, characterized in that: The contrast network includes an encoder and a projection head; The training process of the contrast network includes: Extracting depth features of the base distorted image, the positive sample and the negative sample respectively through an encoder, and extracting respective projection features from the respective depth features according to a projection head; Calculating a first similarity of the projection features between the positive sample pairs, and calculating a second similarity of the projection features between the negative sample pairs; Calculating a contrast loss based on the first similarity and the second similarity; The encoder and the projection head are back-propagated based on the contrastive loss until training is completed.
3. The image quality assessment method according to claim 2, characterized in that: The calculating the contrast loss based on the first similarity and the second similarity includes: ; in, is the contrast loss, is the natural exponential function, is the projection feature of the base distorted image in the positive sample pair And the projection features of the positive sample The first similarity between is the projection feature of the base distorted image in the negative sample pair And the projection features of negative samples The second similarity between is the temperature parameter.
4. The image quality assessment method according to claim 3, characterized in that: The deep network includes a global feature extraction network and a local feature extraction network; The step of extracting the second feature of the image to be evaluated according to the deep network includes: Extracting global features of the image to be evaluated according to the global feature extraction network; Extracting local features of the image to be evaluated according to the local feature extraction network; The second feature of the image to be evaluated is obtained by fusing the global feature and the local feature.
5. The image quality assessment method according to claim 4, characterized in that: The global feature extraction network is a Transformer network, and the local feature extraction network is a CNN network.
6. The image quality assessment method according to claim 5, characterized in that: The evaluating the quality of the image to be evaluated based on the fusion feature includes: Regressing the fused features through a global average pooling layer to obtain the quality of the image to be evaluated; The loss function during the training of the contrast network, the deep network and the global average pooling layer is: ; in, is the number of base distorted images, Output of the global average pooling layer The quality of the base distorted image, For the The actual quality of the base distorted image.
7. The image quality assessment method according to claim 1, characterized in that: The step of fusing the first feature and the second feature of the image to be evaluated by an adaptive weighted attention mechanism to obtain a fused feature includes: Mapping the first feature of the image to be evaluated into a query vector; Mapping the second feature of the image to be evaluated into a key and value vector; The fused feature is calculated using an adaptive weighted attention mechanism formula according to the key and value vectors and the query vector.
8. An image quality assessment system, characterized in that: The system comprises: An image acquisition module, used for acquiring an image to be evaluated; A first calculation module is used to extract a first feature of the image to be evaluated according to a contrast network; wherein the contrast network is obtained by contrast training between positive sample pairs and negative sample pairs based on training data, the training data includes a basic distorted image and positive samples and negative samples obtained by transformation based on the basic distorted image, the positive sample has the same distortion type as the basic distorted image, the negative sample has a different distortion type from the basic distorted image, the positive sample pair is composed of the basic distorted image and the positive sample, and the negative sample pair is composed of the basic distorted image and the negative sample; A second computing module, used for extracting a second feature of the image to be evaluated according to a deep network; A feature fusion module, used for fusing the first feature and the second feature of the image to be evaluated by an adaptive weighted attention mechanism to obtain a fused feature; The quality assessment module is used to assess the quality of the image to be assessed based on the fusion feature.
9. An electronic device, characterized in that: include: at least one controller and a memory for communicatively coupling with the at least one controller; The memory stores instructions that can be executed by the at least one controller, and the instructions are executed by the controller to enable the controller to perform the image quality assessment method according to any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer-executable instructions, and the computer-executable instructions are used to enable a computer to execute the image quality assessment method according to any one of claims 1 to 7.
Citation Information
Patent Citations
No-reference image quality evaluation method based on self-supervised learning and Transform
CN116029953A
No-reference image quality evaluation method based on knowledge distillation and comparative learning
CN116912217A
Intelligent water meter image quality automatic evaluation method based on deep learning
CN118941927A
Image quality evaluation continuous learning method based on prototype playback strategy
CN119723255A
Perception quality evaluation of an object detection system
WO2024186651A1