A method, apparatus, electronic device, and storage medium for evaluating image quality.

By encoding and decoding the image features and quality features of the target image, generating a reconstructed image and calculating the difference value, the problem of high labor intensity and inconsistent evaluation caused by manual reliance in the prior art is solved, and efficient and automated image quality evaluation is achieved.

CN116012341BActive Publication Date: 2026-01-30ZHEJIANG DAHUA TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310035685.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-01-10
Publication Date
2026-01-30
Estimated Expiration
2043-01-10

AI Technical Summary

Technical Problem

In existing technologies, image quality evaluation relies heavily on manual labor, resulting in high labor intensity, low efficiency, and inability to guarantee consistency.

Method used

By acquiring the image features and quality features of the target image, encoding and decoding processes are performed respectively to generate a first reconstructed image and a second reconstructed image, and the feature difference value between the two is calculated to evaluate the image quality.

Benefits of technology

It achieves automated, efficient, and consistent image quality evaluation, reduces manual intervention, and improves the recognition performance of the recognition system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116012341B_ABST
    Figure CN116012341B_ABST
Patent Text Reader

Abstract

This application relates to the field of image processing technology, specifically to an image quality evaluation method, apparatus, electronic device, and storage medium for adaptive image quality evaluation. The method involves acquiring image features obtained by encoding a target image, and acquiring quality features of the target object in the target image extracted based on a task recognition model; performing target decoding processing on the image features to generate a first reconstructed image, and performing target decoding processing on the quality features to generate a second reconstructed image; calculating the image feature difference value between the first and second reconstructed images to obtain the quality evaluation information of the target image. Based on the above method, the problems of high labor intensity, low efficiency, and inability to guarantee consistent image quality evaluation caused by the heavy reliance on manual labor in related technologies are addressed, demonstrating reasonable quality evaluation and robustness.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of image processing technology, and in particular to a method, apparatus, electronic device, and storage medium for evaluating image quality. Background Technology

[0002] In business scenarios involving image recognition tasks, recognition systems are often affected by unconstrained environments, resulting in poor accuracy and stability. To address this, we propose a method for evaluating image quality. By calculating image quality assessment information, images with high evaluation scores are retained, while those with low scores are discarded, enabling the recognition system to perform better image recognition and analysis.

[0003] However, in related technologies, the evaluation of image quality relies on manual selection. For example, in order to improve the recognition performance of the recognition system for pedestrian recognition tasks, it is necessary to manually select the target area where the target object is located and other areas in the image in order to calculate the quality assessment information of the image based on the target area and other areas.

[0004] It can be seen that the relevant technologies are highly dependent on manual labor, which is labor-intensive, inefficient, and influenced by human experience, resulting in large fluctuations in the selection of target areas and making it impossible to guarantee the consistency of image quality evaluation. Summary of the Invention

[0005] This application provides a method, apparatus, electronic device, and storage medium for evaluating image quality, in order to solve the problems of high labor intensity, low efficiency, and inability to guarantee the consistency of image quality evaluation caused by the high dependence on manual labor in related technologies.

[0006] In a first aspect, this application provides a method for evaluating image quality, the method comprising:

[0007] The image features obtained by encoding the target image are acquired, and the quality features of the target object in the target image are acquired based on the task recognition model.

[0008] The image features are subjected to target decoding processing to generate a first reconstructed image, and the quality features are subjected to the same target decoding processing to generate a second reconstructed image;

[0009] The image feature difference values ​​between the first reconstructed image and the second reconstructed image are calculated to obtain the quality assessment information of the target image.

[0010] As one possible implementation, the task identification model includes at least one candidate quality identification model, then,

[0011] The step of obtaining the quality features of the target object in the target image extracted based on the task recognition model includes:

[0012] Based on the target task, a target quality recognition model for recognizing target objects in the target image is determined from the candidate quality recognition models in the task recognition model;

[0013] The target quality recognition model is used to identify target objects in the target image and generate the quality features.

[0014] As one possible implementation, obtaining the image features obtained by encoding the target image includes:

[0015] Based on a preset encoder, the target image is downsampled to obtain the image features.

[0016] As one possible implementation, if the feature dimension of the image feature is a specified dimension, then...

[0017] The step of using the target quality recognition model to identify target objects in the target image and generate the quality features includes:

[0018] The quality recognition model with added global pooling layer is used to embed the quality features of the target object in the target image to obtain the embedding result.

[0019] The embedding result is normalized so that the quality features extracted by the embedding process are in the specified dimension, thus obtaining the quality features.

[0020] As one possible implementation, the quality feature and the image feature have the same feature dimension.

[0021] As one possible implementation, calculating the image feature difference value between the first reconstructed image and the second reconstructed image to obtain the quality assessment information of the target image includes:

[0022] Determine a first feature distribution of the first reconstructed image on the hyperplane, and determine a second feature distribution of the second reconstructed image on the hyperplane;

[0023] The quality assessment information of the target image is obtained based on the first feature distribution and the second feature distribution.

[0024] As one possible implementation, obtaining the quality assessment information of the target image based on the first feature distribution and the second feature distribution includes:

[0025] The distribution difference between the first feature distribution and the second feature distribution is measured, and the distribution difference value is used as the quality assessment information of the target image.

[0026] Secondly, this application provides an image quality evaluation device, the device comprising:

[0027] The acquisition module acquires image features obtained by encoding the target image, and acquires quality features of the target object in the target image extracted based on the task recognition model;

[0028] The generation module performs target decoding processing on the image features to generate a first reconstructed image, and performs the same target decoding processing on the quality features to generate a second reconstructed image.

[0029] The calculation module calculates the image feature difference values ​​between the first reconstructed image and the second reconstructed image to obtain the quality assessment information of the target image.

[0030] As one possible implementation, the task identification model includes at least one candidate quality identification model, then,

[0031] The acquisition module is used to acquire the quality features of the target object in the target image extracted based on the task recognition model, specifically for:

[0032] Based on the target task, a target quality recognition model for recognizing target objects in the target image is determined from the candidate quality recognition models in the task recognition model;

[0033] The target quality recognition model is used to identify target objects in the target image and generate the quality features.

[0034] As one possible implementation, the acquisition module is used to acquire image features obtained by encoding the target image, specifically for:

[0035] Based on a preset encoder, the target image is downsampled to obtain the image features.

[0036] As one possible implementation, if the feature dimension of the image feature is a specified dimension, then...

[0037] The acquisition module is used to identify target objects in the target image using the target quality recognition model and generate the quality features, specifically for:

[0038] The quality recognition model with added global pooling layer is used to embed the quality features of the target object in the target image to obtain the embedding result.

[0039] The embedding result is normalized so that the quality features extracted by the embedding process are in the specified dimension, thus obtaining the quality features.

[0040] As one possible implementation, the quality feature and the image feature have the same feature dimension.

[0041] As one possible implementation, the computing module is specifically used for:

[0042] Determine a first feature distribution of the first reconstructed image on the hyperplane, and determine a second feature distribution of the second reconstructed image on the hyperplane;

[0043] The quality assessment information of the target image is obtained based on the first feature distribution and the second feature distribution.

[0044] As one possible implementation, the calculation module is used to obtain quality assessment information of the target image based on the first feature distribution and the second feature distribution, specifically for:

[0045] The distribution difference between the first feature distribution and the second feature distribution is measured, and the distribution difference value is used as the quality assessment information of the target image.

[0046] Thirdly, this application provides an electronic device, the electronic device comprising:

[0047] Memory, used to store computer programs;

[0048] When a processor executes a computer program stored in the memory, it implements the steps of the image quality evaluation method described above.

[0049] Fourthly, this application provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the above-described image quality evaluation method steps.

[0050] Fifthly, a computer program product comprising: computer program code, which, when run on a computer, causes the computer to perform the steps of the image quality evaluation method described above.

[0051] In this embodiment, image features and quality features of the target image are extracted in different ways, and then the image features and quality features are decoded respectively to generate a first reconstructed image and a second reconstructed image. The difference in image features between the first reconstructed image and the second reconstructed image is used as the quality assessment information of the target image. This can solve the problems of high labor intensity, low efficiency and inability to guarantee the consistency of image quality evaluation caused by the high dependence on manual labor in related technologies. In addition, a task recognition model is introduced here. The image quality of the target image is evaluated by the reproducibility of the measurement features such as extracting quality features and decoding quality features to generate a second reconstructed image. This has reasonable quality evaluation and robustness.

[0052] Other features and advantages of this application will be set forth in the description which follows, and will be apparent in part from the description, or may be learned by practicing the application. The objectives and other advantages of this application may be realized and obtained by means of the structures particularly pointed out in the written description, claims, and drawings. Attached Figure Description

[0053] Figure 1 A schematic diagram illustrating a possible implementation environment for this application;

[0054] Figure 2 A flowchart of an image quality evaluation method provided in this application;

[0055] Figure 3 This application provides a schematic diagram of the structure of a self-attention autoencoder;

[0056] Figure 4 A schematic diagram of the structure of a task feature generator provided in this application;

[0057] Figure 5 This application provides a schematic diagram of calculating quality assessment information based on feature distribution.

[0058] Figure 6 A schematic diagram of a system framework for evaluating image quality provided in this application;

[0059] Figure 7 A schematic diagram of an image quality evaluation device provided in this application;

[0060] Figure 8 A schematic diagram of the structure of an electronic device provided in this application. Detailed Implementation

[0061] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings.

[0062] In this embodiment of the application, the collection, storage, use, processing, transmission, provision, and disclosure of user personal information comply with relevant laws and regulations and do not violate public order and good morals.

[0063] First, some terms used in the embodiments of this application will be explained to facilitate understanding by those skilled in the art.

[0064] Embedding: A distributed representation method used to solve the representation problem under high-dimensional features. Its essence is "compression", that is, using lower-dimensional k (k≥1) features to describe higher-dimensional n (n≥k) features with redundant information.

[0065] For example, in the embodiments of this application, the quality recognition model can be used to embed the quality features of the target object in the target image. The embedding result is the quality features extracted after removing redundant information about the quality of the target object in the target image.

[0066] Image representation: In machine learning, repeatability generally refers to technical repeatability, which mainly focuses on whether the results can be accurately reproduced under the same technical conditions.

[0067] It should be noted that this solution can be applied to image quality evaluation in one or more application scenarios such as security management, intelligent monitoring, smart cities, and pedestrian recognition. This solution is also suitable for tasks requiring adaptive image quality evaluation and improved recognition performance.

[0068] The implementation entity of this solution can be a terminal device, server, or other computing device. By deploying it on relevant computing devices and calculating the difference between image features and quality features extracted from the target image in different ways as quality assessment information, the image quality of the target image is evaluated to improve the recognition performance of the recognition system. Of course, this is only an illustrative example of the entities on which this solution can be applied, and is not specifically limited thereto.

[0069] The terminal device refers to an electronic device with computing capabilities. This electronic device can be a smart camera, personal computer, mobile phone, tablet computer, laptop, e-book reader, smart home device, or other computer device with certain computing capabilities that runs instant messaging software and websites. The server can be a standalone physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, CDN (Content Delivery Network), and big data and artificial intelligence platforms.

[0070] In this embodiment, the task recognition model and the preset encoder can be deployed on a computing device. The task recognition model is used to extract quality features, and the preset encoder is used to extract image features and generate a reconstructed image. The task recognition model may include multiple quality recognition models, and different quality recognition models are suitable for different recognition tasks. For example, the quality recognition model corresponding to the pedestrian recognition task can be used to extract quality features from the target image. The preset encoder can be divided into a preset encoder and a preset decoder. The preset encoder is used to encode the target image to obtain its image features, and the preset decoder is used to decode the quality features and image features respectively to generate a first reconstructed image and a second reconstructed image.

[0071] In one possible design, to facilitate reduced communication latency or load balancing, the execution entity is a server deployed in various regions. Multiple servers share data via blockchain, and these servers essentially constitute a data-sharing system composed of multiple services.

[0072] Each server in the data sharing system has a corresponding node identifier. Each server can also store node identifiers of other servers in the system, enabling subsequent broadcasting of generated blocks to other servers based on their node identifiers. Each server can maintain a node identifier list, storing the server name and node identifier in this list. The node identifier can be an Internet Protocol (IP) address or any other information that can be used to identify the node.

[0073] The design concept of the image quality evaluation method provided in the embodiments of this application will be briefly introduced below.

[0074] With the rapid development of information technology, recognition systems are widely used in daily life. As an example, in the business scenario of pedestrian recognition tasks, pedestrian images are frequently captured, stored, recognized, and analyzed in numerous practical application systems. For instance, body weight recognition, pedestrian attribute recognition, and face recognition are widely used for identity authentication in video surveillance and financial computer vision applications. Due to the influence of various unconstrained environments in reality, the recognition accuracy of pedestrian images can significantly decrease, leading to unstable recognition results. Furthermore, various other factors during the acquisition, compression, transmission, and storage of pedestrian images can also cause a decline in recognition performance.

[0075] To address the aforementioned issues, enhance the generalized recognition capabilities of the recognition system to improve its accuracy and ensure recognition stability, a method for evaluating image quality is proposed. This method calculates image quality assessment information, retains images with high evaluation scores, and discards images with low evaluation scores, thereby enabling the recognition system to better recognize and analyze images.

[0076] In related technologies, image quality evaluation typically relies on manual selection. For example, to improve the recognition performance of a system for pedestrian identification tasks, for target images containing objects, manual selection of the target area and other regions is required. Only then, based on the first depth map information of the target area and the second depth map information of the other regions, can the quality assessment information of the target image be calculated. In this approach, if errors occur during region selection, the final calculated quality assessment information will be inaccurate. Furthermore, the quality assessment information obtained based on the first and second depth map information cannot determine whether the target object in the image is occluded. In scenarios where the target object is severely occluded, this method will seriously affect subsequent extraction of effective information from the target image, leading to inaccurate final quality assessment information.

[0077] In view of this, this application proposes an image quality evaluation method, apparatus, electronic device, and storage medium. The method involves acquiring image features obtained by encoding a target image, and acquiring quality features of a target object in the target image extracted based on a task recognition model; performing target decoding processing on the image features to generate a first reconstructed image, and performing target decoding processing on the quality features to generate a second reconstructed image; calculating the image feature difference value between the first and second reconstructed images to obtain quality evaluation information of the target image.

[0078] In this embodiment, image features and quality features of the target image are extracted in different ways, and then the image features and quality features are decoded respectively to generate a first reconstructed image and a second reconstructed image. The difference in image features between the first reconstructed image and the second reconstructed image is used as the quality assessment information of the target image. This can solve the problems of high labor intensity, low efficiency and inability to guarantee the consistency of image quality evaluation caused by the high dependence on manual labor in related technologies. In addition, a task recognition model is introduced here. The image quality of the target image is evaluated by the reproducibility of the measurement features such as extracting quality features and decoding quality features to generate a second reconstructed image. This has reasonable quality evaluation and robustness.

[0079] The image quality evaluation method provided in this application embodiment can be applied to... Figure 1 The implementation environment shown may include at least camera nodes, operation nodes, management nodes, computing nodes, and storage nodes.

[0080] The camera nodes include spherical and cylindrical cameras, which can be deployed at urban intersections, etc., for capturing images or recording videos. This embodiment focuses on a single target image; the processing of each image in the image sequence dataset is similar. As an example of a business scenario, in a pedestrian recognition task, the target image captured by the camera node at least contains the target object, such as a pedestrian. Of course, in other recognition tasks, the target object can also be other objects such as vehicles.

[0081] The operation node is used to interact with the user, enabling the user to deploy, configure, and manage image quality assessment tasks. For example, in this embodiment, the user can specify a quality recognition model from the task recognition model to extract quality features from the target image, based on the target image recognition task.

[0082] The management node is used to obtain target images or image sequence datasets from the camera nodes, for example, see [link to documentation]. Figure 1 The camera nodes upload the target images or image sequence datasets to the cloud (using cloud storage technology), and the management nodes retrieve the target images or image sequence datasets from the cloud. The management nodes are also used to manage the computing nodes and storage nodes in conjunction with image quality assessment tasks. During the management process, the management nodes forward the target images or image sequence datasets to the computing nodes.

[0083] The computing nodes are used to complete the computational tasks involved in the image quality assessment task based on the received target image or image sequence dataset. This includes decoding the image features and quality features of a single image, and calculating the image feature difference values ​​between the first and second reconstructed images generated from the decoding to obtain the quality assessment information for each image. The same processing is performed on each image in the image sequence dataset, and the calculation result for the image sequence dataset is the quality assessment information corresponding to each image; that is, the calculation result is the corresponding quality assessment information sequence.

[0084] The storage node is used to store target images captured by camera nodes, recorded image sequence datasets, and quality assessment information sequences of target images and image sequence datasets generated through image quality assessment tasks, in order to trace the source.

[0085] In one possible implementation environment, storage nodes can use cloud storage technology for storage. Cloud storage is a new concept that extends and develops from the concept of cloud computing. A distributed cloud storage system (hereinafter referred to as a storage system) refers to a storage system that uses cluster applications, grid technology, and distributed storage file systems to aggregate a large number of storage devices of various types in the network to work together through application software or application interfaces to jointly provide data storage and business access functions.

[0086] It should be noted that the camera node, management node, computing node, and storage node mentioned above are different devices, or any two or three of these nodes can be integrated into the same device. The camera node and storage node are not mandatory, and this solution does not impose specific limitations on them.

[0087] The image quality evaluation method provided in the embodiments of this application will be described in detail below. A single target image is used as an example here; those skilled in the art will understand that the same principle applies to target image sequence datasets. (Reference) Figure 2 The method includes steps 201-203, as detailed below.

[0088] Step 201: Obtain the image features obtained by encoding the target image, and obtain the quality features of the target object in the target image extracted based on the task recognition model;

[0089] In this embodiment, the target image is downsampled based on a preset encoder to obtain image features. The task recognition model includes at least one candidate quality recognition model. From the task recognition models, a target quality recognition model is determined based on the target task to identify target objects in the target image. Then, the target quality recognition model is used to identify the target objects in the target image, generating quality features of a specified dimension.

[0090] The aforementioned target task is a task designed to meet business requirements. The target task indicates the target object to be identified. Based on the target task, a target quality identification model can be determined to identify the target object. Here, the target task and / or target quality identification model can be obtained by accepting user instructions, or the target quality identification model can be determined based on a preset correspondence between the target task / target object and each candidate quality identification model.

[0091] As an optional implementation, in order to better obtain the image representation of the target image and have the ability to reconstruct and embed, the preset encoder is designed as a self-attention autoencoder, which consists of a preset encoder and a preset decoder.

[0092] like Figure 3As shown, this is a self-attention autoencoder. In the context of pedestrian recognition, the convolutional neural network can be ResNet34 (a 34-layer deep residual network), which is used to extract preliminary features of the target image.

[0093] It should be noted that two multi-head self-attention transducer (MHSA) blocks can be added at appropriate locations in the self-attention autoencoder to effectively improve the representativeness of the attention region (the target region where the target object is located) in the image.

[0094] In one application example of the preset encoder, preliminary features of the target image are extracted based on ResNet34. The features are then enhanced by a multi-head self-attention converter module to improve the representativeness of the attention region. After downsampling the preliminary image features, they are passed through the multi-head self-attention converter module again to further enhance the representativeness of the attention region. The output is then input into a one-dimensional global average pooling (GAP) layer to stretch the feature length and maintain the similarity of the feature data distribution, thus obtaining the image features of the target image.

[0095] In one application example of the preset decoder, the image features of the target image are upsampled to complete the decoding process and obtain a reconstructed image generated based on the image features.

[0096] Furthermore, for training the adaptive autoencoder, given a training set X of length n, a dataset D consisting of paired images, D = {(x1,x1),(x2,x2),...,(x...}, can be used. n ,x n )|x i The image is generated by training on the target image (∈X). Here, the preset encoder in the generated adaptive autoencoder can extract image features. For example, in this embodiment, the target image is input into the preset encoder, and the output is obtained as follows:

[0097] E SA =W en (I)

[0098] Among them, E SA Let I be the target image, and W be the image features extracted by the preset encoder. en This is the preset encoder.

[0099] It should be noted that, regardless of the business scenario, the dimension of the image features extracted by the preset encoder is the specified dimension. In this embodiment, the image features with values ​​between 0 and 1 are used as an example.

[0100] As an optional implementation, since the definition of image quality differs in different recognition tasks, this approach utilizes a task-based recognition model to accommodate as many quality recognition models as possible for different recognition tasks, with a correspondence between the quality recognition models and the recognition tasks. For example, in the context of pedestrian recognition, the pedestrian quality recognition model, as the backbone of the model, can be deployed in scenarios such as... Figure 4 In the task feature generator shown, the target task in this embodiment is a specified identification task.

[0101] In one application example, each quality recognition model in the task recognition model can determine a target quality recognition model based on the target task. Based on the target quality recognition model, preliminary quality features are extracted from the input target image. The output preliminary quality features are then flattened into task embedding, and a one-dimensional global pooling layer is added to ensure that the output lengths of different quality recognition models in the task recognition model are the same, thereby reducing the differences caused by the training settings of different quality recognition models in the task recognition model. Then, the features are embedded to obtain the quality features of the target image.

[0102] Furthermore, since different quality recognition models in the task recognition model can have different parameter settings during pre-training, the values ​​embedded in the output of different quality recognition models vary significantly. Directly comparing the output value of this quality feature with the aforementioned image features is unreasonable. For example, the image feature values ​​generated by the target image based on the preset encoder are between 0 and 1, while the quality feature values ​​embedded by the task recognition model are between 0 and 100. These significantly different values ​​cannot be directly compared. To address this, this application employs a normalization process to normalize the values ​​of the embedded quality features, generating quality features with the same specified dimensions as the image features, thereby eliminating the aforementioned influence.

[0103] Based on the above operations, the generated quality features for the specified dimension can be seen as follows:

[0104] E Task =norm(GAP(W task (I)))

[0105] Among them, E Task For quality features of a specified dimension, norm() performs normalization, GAP() is the output of a one-dimensional global pooling layer, and W... Task Here, I represents the quality recognition model in the task recognition model, and I is the target image.

[0106] It should be noted that normalization is a preferred implementation method proposed in the embodiments of this application. As long as the feature dimensions of the quality features and the image features are the same, no specific restrictions are imposed here.

[0107] The above implementation methods can obtain image features obtained by encoding the target image, as well as quality features of the target object in the target image extracted based on the task recognition model.

[0108] Step 202: Perform target decoding processing on the image features to generate a first reconstructed image, and perform the target decoding processing on the quality features to generate a second reconstructed image;

[0109] In this embodiment of the application, in order to enhance the feature difference between image features and quality features, the image features and quality features are decoded based on a preset decoder to obtain a first reconstructed image generated by image feature reconstruction and a second reconstructed image generated by quality feature reconstruction.

[0110] Specifically, the above process can be seen as follows:

[0111] I SA =W de (E SA )

[0112] I Task =W de (E Task )

[0113] Among them, I SA For the first reconstructed image generated from image features, I Task W is the second reconstructed image generated from the quality features. de E is the preset decoder. SA For image features, E Task For quality characteristics.

[0114] It should be noted that the description of the above-mentioned preset decoder can be found in the self-attention autoencoder, and will not be repeated here.

[0115] Step 203: Calculate the image feature difference value between the first reconstructed image and the second reconstructed image to obtain the quality assessment information of the target image.

[0116] In this embodiment, the quality assessment information of the target image is measured by comparing the difference values ​​of the first reconstructed image and the second reconstructed image on the hyperplane.

[0117] As one implementation method, such as Figure 5The diagram illustrates the calculation of quality assessment information based on feature distribution. First, the first feature distribution of the first reconstructed image on the hyperplane is determined, and the second feature distribution of the second reconstructed image on the hyperplane is determined. Then, the distribution difference value between the first feature distribution and the second feature distribution is measured, and the distribution difference value is used as the quality assessment information of the target image.

[0118] Specifically, the above process can be seen as follows:

[0119] Q(I)=F(I SA ,I Task )

[0120] Among them, I SA For the first reconstructed image of the target image, I Task For the second reconstructed image of the target image, F() is I SA and I Task The mapping function to Q(I) is where Q() is the distribution difference value of the first reconstructed image and the second reconstructed image on the hyperplane (i.e., the quality assessment information of the target image).

[0121] It should be noted that the above measurement operation adopts a distribution difference measurement method. For example, the optimal transmission distance (Wasserstein Distance) can be used to evaluate the distribution difference value of the first reconstructed image and the second reconstructed image on the hyperplane, so as to obtain the quality assessment information of the target image.

[0122] In addition, the aforementioned quality assessment information may be a distribution difference value or a quality assessment score calculated by other methods. Of course, the quality assessment information may also include other information used to evaluate image quality, but this application embodiment does not specifically limit this.

[0123] In accordance with the above implementation method, to measure the recognition performance of the recognition system, the first reconstructed image is used as the image representation of the target image. A second reconstructed image is generated based on the quality recognition model corresponding to the target task in the task recognition model, which serves as the evaluation of the image quality of the target image by the target task. Then: when the target task considers the quality of the target image to be poor, its extracted quality features are blurry and stable, making it difficult to restore the image. Therefore, the distribution of the second reconstructed image generated by the decoding process differs greatly from that of the first reconstructed image on the hyperplane. When the target task considers the quality of the target image to be good, its extracted quality features are robust and stable, and the restored image is closer to the original image. Therefore, the difference between the second reconstructed image generated by the decoding process and the first reconstructed image on the hyperplane is smaller.

[0124] Better place, Figure 6This is a schematic diagram of a system framework for evaluating image quality provided in an embodiment of this application. The target image is input to a task feature generator module and a self-attention autoencoder module. The task feature generator extracts quality features from the target image using a quality recognition model corresponding to the target task and generates a first reconstructed image through an upsampling operation. The self-attention autoencoder obtains image features of the target image based on a preset encoder and generates a second reconstructed image through an upsampling operation. The difference in distribution between the first and second reconstructed images on the hyperplane is measured to obtain quality assessment information for the target image.

[0125] In summary, in the embodiments of this application, image features and quality features of the target image are extracted in different ways, and then the image features and quality features are decoded respectively to generate a first reconstructed image and a second reconstructed image. The difference value of image features between the first reconstructed image and the second reconstructed image is used as the quality evaluation information of the target image. The image quality can be adaptively evaluated according to the target task, and it has reasonable quality evaluation and robustness.

[0126] Based on the embodiments of this application, at least the following technical effects can be achieved:

[0127] To address the problem that recognition systems cannot adaptively predict image quality based on relevant recognition tasks, an end-to-end task-guided image quality assessment method is proposed. This method achieves real-time point-to-point prediction, rather than being limited to image sequence datasets of a single target object. The system framework provided by this method can be regarded as a general framework for image quality perception, which is versatile and can be applied to other similar tasks, and is not limited to the pedestrian recognition example in the embodiments of this application.

[0128] Based on the same inventive concept, this application also provides an image quality evaluation device to solve the problems of high labor intensity, low efficiency, and inability to guarantee consistent image quality evaluation caused by the high reliance on manual labor in related technologies. See [link to related document]. Figure 7 The device includes:

[0129] The acquisition module 701 acquires image features obtained by encoding the target image, and acquires quality features of the target object in the target image extracted based on the task recognition model;

[0130] The generation module 702 performs target decoding processing on the image features to generate a first reconstructed image, and performs the target decoding processing on the quality features to generate a second reconstructed image;

[0131] The calculation module 703 calculates the image feature difference value between the first reconstructed image and the second reconstructed image to obtain the quality assessment information of the target image.

[0132] As one possible implementation, the task identification model includes at least one candidate quality identification model, then,

[0133] The acquisition module 701 is used to acquire the quality features of the target object in the target image extracted based on the task recognition model, specifically for:

[0134] Based on the target task, a target quality recognition model for recognizing target objects in the target image is determined from the candidate quality recognition models in the task recognition model;

[0135] The target quality recognition model is used to identify target objects in the target image and generate the quality features.

[0136] As one possible implementation, the acquisition module 701 is used to acquire image features obtained by encoding the target image, specifically for:

[0137] Based on a preset encoder, the target image is downsampled to obtain the image features.

[0138] As one possible implementation, if the feature dimension of the image feature is a specified dimension, then...

[0139] The acquisition module 701 is used to identify target objects in the target image using the target quality recognition model and generate the quality features, specifically for:

[0140] The quality recognition model with added global pooling layer is used to embed the quality features of the target object in the target image to obtain the embedding result.

[0141] The embedding result is normalized so that the quality features extracted by the embedding process are in the specified dimension, thus obtaining the quality features.

[0142] As one possible implementation, the quality feature and the image feature have the same feature dimension.

[0143] As one possible implementation, the computing module 703 is specifically used for:

[0144] Determine a first feature distribution of the first reconstructed image on the hyperplane, and determine a second feature distribution of the second reconstructed image on the hyperplane;

[0145] The quality assessment information of the target image is obtained based on the first feature distribution and the second feature distribution.

[0146] As one possible implementation, the calculation module 703 is used to obtain the quality assessment information of the target image based on the first feature distribution and the second feature distribution, specifically for:

[0147] The distribution difference between the first feature distribution and the second feature distribution is measured, and the distribution difference value is used as the quality assessment information of the target image.

[0148] Based on the same inventive concept, this application also provides an electronic device that can realize the function of the aforementioned image quality evaluation device. (Refer to...) Figure 8 The electronic device includes:

[0149] At least one processor 801 and a memory 802 connected to at least one processor 801. In this embodiment, the specific connection medium between the processor 801 and the memory 802 is not limited. Figure 8 The example shown is the connection between processor 801 and memory 802 via bus 800. Bus 800 is... Figure 8 The connections between other components are indicated by thick lines and are for illustrative purposes only, not as limiting information. The 800 bus can be divided into address bus, data bus, control bus, etc., for ease of representation. Figure 8 The term is represented by a single thick line, but this does not imply that there is only one bus or one type of bus. Alternatively, the processor 801 can also be called a controller; there is no restriction on the name.

[0150] In this embodiment, memory 802 stores instructions executable by at least one processor 801. By executing the instructions stored in memory 802, at least one processor 801 can perform the image quality evaluation method discussed above. Processor 801 can implement... Figure 7 The functions of each module in the device / system shown.

[0151] The processor 801 is the control center of the device / system. It can connect to various parts of the control device through various interfaces and lines. By running or executing instructions stored in memory 802 and calling data stored in memory 802, it can monitor the various functions and data processing of the device / system as a whole.

[0152] In one possible design, processor 801 may include one or more processing units. Processor 801 may integrate an application processor and a modem processor, wherein the application processor mainly handles the operating system, user interface, and applications, and the modem processor mainly handles wireless communication. It is understood that the modem processor may also not be integrated into processor 801. In some embodiments, processor 801 and memory 802 may be implemented on the same chip; in some embodiments, they may also be implemented on separate chips.

[0153] The processor 801 can be a general-purpose processor, such as a central processing unit (CPU), digital signal processor, application-specific integrated circuit, field-programmable gate array or other programmable logic device, discrete gate or transistor logic device, or discrete hardware component, capable of implementing or executing the methods, steps, and logic block diagrams disclosed in the embodiments of this application. The general-purpose processor can be a microprocessor or any conventional processor. The steps of the image quality evaluation method disclosed in the embodiments of this application can be directly manifested as being executed by a hardware processor, or executed by a combination of hardware and software modules within the processor.

[0154] Memory 802, as a non-volatile computer-readable storage medium, can be used to store non-volatile software programs, non-volatile computer-executable programs, and modules. Memory 802 may include at least one type of storage medium, such as flash memory, hard disk, multimedia card, card-type memory, random access memory (RAM), static random access memory (SRAM), programmable read-only memory (PROM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), magnetic storage, magnetic disk, optical disk, etc. Memory 802 can be any other medium capable of carrying or storing desired program code in the form of instructions or data structures that can be accessed by a computer, but is not limited thereto. In the embodiments of this application, memory 802 can also be a circuit or any other device / system capable of implementing storage functions for storing program instructions and / or data.

[0155] By designing and programming the processor 801, the code corresponding to the image quality evaluation method described in the foregoing embodiments can be embedded into the chip, enabling the chip to execute it during operation. Figure 2The steps of the image quality evaluation method in the illustrated embodiment are described. How to design and program the processor 801 is a technique well-known to those skilled in the art and will not be elaborated upon here.

[0156] Based on the same inventive concept, embodiments of this application also provide a storage medium storing computer instructions that, when executed on a computer, cause the computer to perform the image quality evaluation method described above.

[0157] In some possible implementations, various aspects of the image quality evaluation method provided in this application can also be implemented as a program product, comprising program code. When the program product is run on a computer device, the program code causes the computer device to perform the steps of the image quality evaluation method according to the various exemplary embodiments of this application described above. For example, the computer device can perform actions such as... Figure 2 The steps are shown in the figure.

[0158] The program product may employ any combination of one or more readable media. A readable medium may be a readable signal medium or a readable storage medium. A readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples (a non-exhaustive list) of readable storage media include: electrical connections having one or more wires, portable disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof.

[0159] The program product of the embodiments of this application may employ a portable compact disc read-only memory (CD-ROM) and include program code, and may run on a computing device. However, the program product of this application is not limited thereto. In the embodiments of this application, the readable storage medium may be any tangible medium that contains or stores a program that may be used by or in conjunction with a command execution system, apparatus, or device.

[0160] A readable signal medium may include a data signal propagated in baseband or as part of a carrier wave, carrying readable program code. This propagated data signal may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A readable signal medium may also be any readable medium other than a readable storage medium, capable of sending, propagating, or transmitting a program for use by or in conjunction with a command execution system, apparatus, or device.

[0161] The program code contained on the readable medium may be transmitted using any suitable medium, including but not limited to wireless, wired, optical fiber, RF, etc., or any suitable combination thereof.

[0162] Program code for performing the operations of this application can be written in any combination of one or more programming languages, including object-oriented programming languages ​​such as Java and C++, and conventional procedural programming languages ​​such as C or similar languages. The program code can execute entirely on the user's computing device, partially on the user's device, as a standalone software package, partially on the user's computing device and partially on a remote computing device, or entirely on a remote computing device or server. In cases involving remote computing devices, the remote computing device can be connected to the user's computing device via any type of network, including a local area network (LAN) or a wide area network (WAN), or it can be connected to an external computing device (e.g., via the Internet using an Internet service provider).

[0163] It should be noted that although several units or sub-units of the device have been mentioned in the detailed description above, this division is merely exemplary and not mandatory. In fact, according to embodiments of this application, the features and functions of two or more units described above can be embodied in one unit. Conversely, the features and functions of one unit described above can be further divided and embodied by multiple units.

[0164] Furthermore, although the operations of the method of this application are described in a specific order in the accompanying drawings, this does not require or imply that these operations must be performed in that specific order, or that all the operations shown must be performed to achieve the desired result. Additionally or alternatively, certain steps may be omitted, multiple steps may be combined into one step, and / or one step may be broken down into multiple steps.

[0165] Those skilled in the art will understand that embodiments of this application can be provided as methods, apparatus / systems, or computer program products. Therefore, this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this application can take the form of a computer program product embodied on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0166] This application is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this application. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart... Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.

[0167] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.

[0168] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.

[0169] Obviously, those skilled in the art can make various modifications and variations to this application without departing from the spirit and scope of this application. Therefore, if such modifications and variations fall within the scope of the claims of this application and their equivalents, this application also intends to include such modifications and variations.

Claims

1. A method of evaluating image quality, characterized by, The method comprises: obtaining image features obtained by encoding a target image; determining a target quality recognition model for recognizing a target object in the target image from candidate quality recognition models in a task recognition model, wherein the task recognition model comprises at least one candidate quality recognition model; using the target quality recognition model to recognize the target object in the target image and generate quality features of the target object in the target image; performing target decoding processing on the image features to generate a first reconstructed image, and performing the target decoding processing on the quality features to generate a second reconstructed image; determining a first feature distribution of the first reconstructed image on a hyperplane and a second feature distribution of the second reconstructed image on the hyperplane; measuring a distribution difference value between the first feature distribution and the second feature distribution, and taking the distribution difference value as quality evaluation information of the target image.

2. The method of claim 1, wherein, The method comprises: performing down-sampling processing on the target image based on a preset encoder to obtain the image features.

3. The method of claim 1, wherein, The feature dimension of the image features is a specified dimension, and then The method comprises: using the target quality recognition model with a global pooling layer to perform embedding processing on the quality features of the target object in the target image to obtain an embedding result; performing normalization processing on the embedding result to make the quality features extracted by the embedding processing be in the specified dimension to obtain the quality features.

4. The method according to any one of claims 1 to 3, characterized in that, The feature dimension of the quality features is the same as that of the image features.

5. An image quality evaluation device, characterized by comprising: The device comprises: an obtaining module, which obtains image features obtained by encoding a target image, determines a target quality recognition model for recognizing a target object in the target image from candidate quality recognition models in a task recognition model, wherein the task recognition model comprises at least one candidate quality recognition model, uses the target quality recognition model to recognize the target object in the target image and generate quality features of the target object in the target image, and a generating module, which performs target decoding processing on the image features to generate a first reconstructed image, and performs the target decoding processing on the quality features to generate a second reconstructed image; a calculating module, which determines a first feature distribution of the first reconstructed image on a hyperplane and a second feature distribution of the second reconstructed image on the hyperplane, measures a distribution difference value between the first feature distribution and the second feature distribution, and takes the distribution difference value as quality evaluation information of the target image.

6. An electronic device, comprising: comprise: a memory for storing a computer program; a processor for executing the computer program stored in the memory to implement the method steps in any one of claims 1-4.

7. A computer readable storage medium characterized in that, The computer readable storage medium stores a computer program, and the computer program is executed by a processor to implement the method steps in any one of claims 1-4.

8. A computer program product, characterised in that, The computer program product comprises computer program code, and when the computer program code is executed on a computer, the computer program code causes the computer to perform the method in any one of claims 1-4.

Citation Information

Patent Citations

  • Image evaluation method, image capturing method and digital camera thereof

    CN101800852A

  • Endoscope image detection method, device, storage medium and electronic equipment

    CN113487608A