Image quality determination method, device, terminal and storage medium
By preprocessing the original image and training of twin network models, the scene dependence problem of deep learning methods on image quality determination in the prior art is solved, and image quality determination in different scenarios is achieved, and accuracy and generalization ability are improved.
Patent Information
- Application Number
- CN202210345355.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-03-31
- Publication Date
- 2025-05-09
- Estimated Expiration
- 2042-03-31
AI Technical Summary
In the prior art, deep learning methods depend on the determination of image quality and cannot be applied to the determination of image quality in other scenarios.
By receiving the original image, preprocessing is performed to obtain multiple sub-images, the initial twin network model is trained using the first scene image data set and the second scene image data set to determine the target twin network model, and finally, the image quality of the original image is determined based on the multiple sub-images and the target twin network model.
It realizes the effective determination of image quality in different scenarios, enhances the model's generalization ability of image data in different scenarios, supports image quality detection of different sizes, and improves the accuracy of image quality determination.
Smart Images

Figure CN114926674B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of computer vision technology, and in particular to a method, device, terminal and storage medium for determining image quality. Background Art
[0002] Image quality detection is a basic task in the field of computer vision. Its purpose is to evaluate the quality of images (such as Figure 1 Generally, the image quality is classified, and after the quality type of the image is determined, the image quality detection is implemented based on the image quality type.
[0003] At present, deep learning methods are mainly used to determine image quality, and deep learning methods are mainly divided into score-based, rank-based and multi-task. Among them, the score-based method is the simplest method, and its method structure is as follows: Figure 2 As shown in , the network outputs an image quality score ranging from 0 to 1, and the image quality is determined by the score. Rank-based is transformed from direct score regression to learning the quality ranking of images, such as Figure 3 As shown in the figure, by artificially simulating the distortion type, a two-stage training method is used: training quality ranking -> regression quality score, and an image quality score is obtained to determine the image quality. Multi-task is expected to increase the additional supervision information of image quality through multi-task learning to assist in quality assessment.
[0004] However, deep learning methods are dependent on the scenarios in which the model is trained, which makes it impossible to simultaneously determine the quality of images in other scenarios. Summary of the invention
[0005] The main purpose of the present application is to provide a method, device, terminal and storage medium for determining image quality, so as to solve the problem that the related art cannot be applied to the determination of image quality in other scenarios.
[0006] In order to achieve the above objectives, in a first aspect, the present application provides a method for determining image quality, comprising:
[0007] receiving the original image;
[0008] Preprocess the original image to obtain multiple sub-images;
[0009] Using a first scene image data set and a second scene image data set to train an initial twin network model, and determine a target twin network model, wherein the first scene image data set includes image quality data corresponding to the first scene image;
[0010] Based on multiple sub-images and the target Siamese network model, the image quality of the original image is determined.
[0011] In a possible implementation, the original image is preprocessed to obtain multiple sub-images, including:
[0012] Convert the format of the original image to obtain the original image after format conversion;
[0013] The original image after format conversion is subjected to sliding window processing using a preset sliding window to obtain multiple sub-images.
[0014] In a possible implementation, the initial twin network model is trained using the first scene image data set and the second scene image data set to determine the target twin network model, including:
[0015] Fusing the first scene image dataset and the second scene image dataset to obtain a fused dataset;
[0016] Extracting data from the fused data set to obtain a first fused data set and a second fused data set, wherein the amount of first scene image data included in the first fused data set is greater than the amount of data in the second scene image data set, and the amount of second scene image data included in the second fused data set is greater than the amount of data in the second scene image data set;
[0017] Inputting the data in the second scene image data set and the first fused data set into the first network in the initial twin network model to obtain a first output result, and simultaneously inputting the data in the second scene image data set and the second fused data set into the second network in the initial twin network model to obtain a second output result;
[0018] Based on the first output result, the second output result and the preset loss function, the target twin network model is determined.
[0019] In a possible implementation, the first output result includes a sub-first output result corresponding to the second scene image dataset and a sub-second output result corresponding to the first fused dataset, and the second output result includes a sub-third output result corresponding to the second scene image dataset and a sub-fourth output result corresponding to the second fused dataset;
[0020] Based on the first output result, the second output result and the preset function, determining the target twin network model includes:
[0021] Determine a fifth sub-output result and a sixth sub-output result based on the first sub-output result and the third sub-output result, respectively;
[0022] Iteratively updating the sub-first output result, the sub-second output result, the sub-third output result, the sub-fourth output result, the sub-fifth output result, and the sub-sixth output result using a preset loss function to obtain an initial fitting result;
[0023] When the initial fitting result meets the preset fitting result, the initial twin network model is used as the target twin network model.
[0024] In a possible implementation, determining the fifth sub-output result and the sixth sub-output result based on the first sub-output result and the third sub-output result, respectively, includes:
[0025] Selecting the image quality types corresponding to the images with the highest scores in the sub-first output result and the sub-third output result respectively, to obtain the first image quality type and the second image quality type;
[0026] taking the first image quality type, the score corresponding to the first image quality type, the label data, and the threshold as a fifth sub-output result;
[0027] The second image quality type, the score corresponding to the second image quality type, the label data and the threshold are used as a sixth sub-output result.
[0028] In a possible implementation, the preset loss function includes a first loss function, a second loss function, a third loss function, a fourth loss function, and a fifth loss function, and the initial fitting result includes a first fitting result, a second fitting result, a third fitting result, a fourth fitting result, and a fifth fitting result;
[0029] The sub-first output result, the sub-second output result, the sub-third output result, the sub-fourth output result, the sub-fifth output result and the sub-sixth output result are iteratively updated using a preset loss function to obtain an initial fitting result, including:
[0030] Inputting the sub-second output result, the first image quality type, and the label data corresponding to the first image quality type into a first loss function to obtain a first fitting result;
[0031] Inputting the sub-third output result, the first image quality type, and the score and threshold corresponding to the first image quality type into a second loss function to obtain a second fitting result;
[0032] Inputting the sub-first output result, the second image quality type, and the score and threshold corresponding to the second image quality type into a third loss function to obtain a third fitting result;
[0033] Inputting the fourth output result, the second image quality type, and the label data corresponding to the second image quality type into a fourth loss function to obtain a fourth fitting result;
[0034] The sub-first output result and the sub-third output result are input into a fifth loss function to obtain a fifth fitting result.
[0035] In one possible implementation, determining the image quality of the original image based on the plurality of sub-images and the target Siamese network model includes:
[0036] Inputting the plurality of sub-images into the first network in the target twin network model to obtain a plurality of target results, wherein the plurality of sub-images correspond one-to-one to the plurality of target results;
[0037] Extracting a score corresponding to an image quality type from each of the multiple target results to obtain a score corresponding to an image quality type corresponding to each of the multiple sub-images;
[0038] Compare the score corresponding to the image quality type corresponding to each sub-image with the preset standard score to determine the image quality type corresponding to each sub-image;
[0039] The image quality of the original image is determined based on the image quality type corresponding to each sub-image.
[0040] In a possible implementation manner, the original image is an RGB image.
[0041] In a second aspect, an embodiment of the present invention provides a device for determining image quality, including:
[0042] An image receiving module, used for receiving an original image;
[0043] An image preprocessing module is used to preprocess the original image to obtain multiple sub-images;
[0044] A model training module, used to train an initial twin network model using a first scene image data set and a second scene image data set to determine a target twin network model, wherein the first scene image data set includes image quality data corresponding to the first scene image;
[0045] The image quality determination module is used to determine the image quality of the original image based on multiple sub-images and the target twin network model.
[0046] In a third aspect, an embodiment of the present invention provides a terminal, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the steps of any of the above methods for determining image quality when executing the computer program.
[0047] In a fourth aspect, an embodiment of the present invention provides a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of any of the above methods for determining image quality are implemented.
[0048] The embodiments of the present invention provide a method, device, terminal and storage medium for determining image quality, including: receiving an original image, then preprocessing the original image to obtain multiple sub-images, then using a first scene image data set and a second scene image data set to train an initial twin network model, determine a target twin network model, and finally determine the image quality of the original image based on the multiple sub-images and the target twin network model. The target twin network model of the present invention includes two network models. The two network models enhance the generalization ability of the two networks for image data in different scenes through mutual supervised learning. In addition, the target twin network model can support quality detection of images of different sizes. When the original image is an image of any scene, the target twin network model can effectively output its image quality type, and the determination of image quality will not be affected by the image size problem, thereby improving the accuracy of image quality determination. BRIEF DESCRIPTION OF THE DRAWINGS
[0049] The drawings constituting a part of this application are used to provide a further understanding of this application, so that other features, purposes and advantages of this application become more obvious. The schematic embodiment drawings and their descriptions of this application are used to explain this application and do not constitute an improper limitation on this application. In the drawings:
[0050] Figure 1 is a diagram describing a concept of image quality provided by an embodiment of the present invention;
[0051] Figure 2 is a flowchart of implementing image quality detection using a score-based method provided by an embodiment of the present invention;
[0052] Figure 3 is a flowchart of implementing image quality detection using a rank-based method provided by an embodiment of the present invention;
[0053] Figure 4 is a flow chart of an implementation method of an image quality determination method provided by an embodiment of the present invention;
[0054] Figure 5 is a schematic diagram of the principle of performing sliding window processing on an image provided by an embodiment of the present invention;
[0055] Figure 6 is a flowchart of implementing network processing for training an initial twin network model provided by an embodiment of the present invention;
[0056] Figure 7 It is a flowchart for implementing the fusion of multiple loss functions in training the initial twin network model provided by an embodiment of the present invention;
[0057] Figure 8is a structural schematic diagram of a device for determining image quality provided by an embodiment of the present invention;
[0058] Fig. 9 is a schematic diagram of a terminal provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0059] In order to make the purpose, technical solution and advantages of the embodiments of the present invention clearer, the technical solution in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0060] The terms "first", "second", "third", "fourth", etc. (if any) in the specification and claims of the present invention and the above drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the numbers used in this way can be interchanged where appropriate, so that the embodiments of the present invention described herein can be implemented in sequences other than those illustrated or described herein.
[0061] It should be understood that in various embodiments of the present invention, the size of the sequence number of each process does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present invention.
[0062] It should be understood that in the present invention, "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions. For example, a process, method, system, product or apparatus comprising a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to these processes, methods, products or apparatuses.
[0063] It should be understood that in the present invention, "plurality" refers to two or more than two. "And / or" is merely a description of the association relationship of associated objects, indicating that three relationships may exist. For example, and / or B can represent: A exists alone, A and B exist at the same time, and B exists alone. The character " / " generally indicates that the objects associated before and after are in an "or" relationship. "Contains A, B and C", "Contains A, B, C" means that A, B, and C are all included, "Contains A, B or C" means that one of A, B, and C is included, and "Contains A, B and / or C" means that any one, any two, or any three of A, B, and C are included.
[0064] It should be understood that in the present invention, "B corresponding to A", "B corresponding to A", "A corresponds to B" or "B corresponds to A" means that B is associated with A and B can be determined based on A. Determining B based on A does not mean determining B based only on A, but B can also be determined based on A and / or other information. A and B match when the similarity between A and B is greater than or equal to a preset threshold.
[0065] Depending on the context, "if" as used herein may be interpreted as "when" or "when" or "in response to determining" or "in response to detecting."
[0066] The technical solution of the present invention is described in detail with specific embodiments below. The following specific embodiments can be combined with each other, and the same or similar concepts or processes may not be described in detail in some embodiments.
[0067] In order to make the purpose, technical solutions and advantages of the present invention more clear, specific embodiments will be described below in conjunction with the accompanying drawings.
[0068] In one embodiment, Figure 4 As shown, a method for determining image quality is provided, comprising the following steps:
[0069] Step S401: receiving an original image;
[0070] Step S402: pre-processing the original image to obtain multiple sub-images.
[0071] For image preprocessing, the present invention first performs format conversion on the original image to obtain the format-converted original image, and then performs sliding window processing on the format-converted original image using a preset sliding window to obtain multiple sub-images.
[0072] The following takes the original image as an RGB image as an example to illustrate the image preprocessing process. Specifically, it is assumed that the input shape of the RGB image is (H, W, 3), where H represents the height of the RGB image, W represents the width of the RGB image, and 3 represents the number of channels of the RGB image is 3. Since the RGB image needs to be input into the target twin network model, it is necessary to convert the RGB image into a format (dimensionality conversion) so that the RGB image after the format conversion meets the format of the input data required by the target twin network model. The present application adopts the ToTensor operation to convert the RGB image with an input shape of (H, W, 3) into an image with a shape of (1, 3, H, W). Since the existing technology for image quality cannot effectively support multiple scales of images, the present invention hopes to try not to scale the format-converted image by resizing, but to crop the image by dividing the patch.
[0073] like Figure 5 As shown in the figure, the converted RGB image is processed by sliding window with the patch_size we set, and the image is divided into multiple patches, that is, multiple sub-images, and then input into the target twin network. In this way, prediction can be made without losing the accuracy of the original image. In addition, in order to simulate our preprocessing method during the training process, the original image is randomly cropped.
[0074] Step S403: Use the first scene image data set and the second scene image data set to train the initial twin network model to determine the target twin network model.
[0075] The first scene image data set includes image quality data corresponding to the first scene image, and the image quality data refers to label data obtained by setting a label on the first scene image.
[0076] Since the existing technologies for image quality are dependent on data and scenarios, the present invention adopts a multi-model fusion module based on a twin network to support the prediction of various scenarios such as general OCR, reducing the overfitting of the model to the training set. In addition, the method adopted by the present invention is not expressed in the form of quality scores, but by manually setting the distortion category (i.e., the type of image quality) and identifying the pre-set category.
[0077] Combine the following Figure 6 The initial twin network model shown illustrates the process of determining the target twin network model:
[0078] First, the first scene image dataset and the second scene image dataset are fused to obtain a fused dataset. In terms of training, the present invention uses a domain adaptive method to divide the dataset into a first scene image dataset (also called a source dataset, Source Set) and a second scene image dataset (also called a target dataset, Target Set). The source dataset is characterized by a large amount of labeled data, and the target dataset is characterized by a small amount of labeled data, or even no labeling. The purpose of this design is that we expect to expand data more quickly and efficiently, and to improve the efficiency of model migration. For example, our source dataset contains a large number of labeled but single-scene pictures, that is, the images in the source dataset are first scene images, but we hope that the trained model can not only make predictions in the same scene as the source dataset, but also in other scenes.
[0079] Therefore, we introduced the target dataset and used the twin network method during training, so that the data of the two datasets passed through the two networks respectively, and the output of the network passed through the mutual supervision module (Mutual Module) to establish a connection between the two datasets. This allows the output model to be expanded to more diverse scenarios. Before training, the datasets from the two domains are processed through MixUp, that is, the source dataset and the target dataset are fused to obtain a fused dataset ( Figure 6 The Mix Module shown in ). The calculation formula of Mixup is as follows:
[0080] I m =α*I s +(1-α)*I t
[0081] Among them, α (alpha) is a random number ranging from 0 to 1, which is usually set manually, such as 0.3, 0.5, etc. I_s is an image in the source dataset, I_t is an image in the target dataset, and I_m is the result of mixup of the two images.
[0082] After the data is fused, the model's generalization ability for the source and target data sets is further enhanced through two models (Net1 and Met2) and a mutual supervision module (MutualModule). Since the first scene images corresponding to the source data set all contain label data, but the second scene images corresponding to the target data set rarely contain label data, in order to achieve the supervision effect of the two network models in the initial twin model, different data are input into the two network models when training the initial twin model to provide conditions for subsequent fitting.
[0083] Based on the difference in input data of different network models, the data in the fused data set must be extracted first to obtain the first fused data set and the second fused data set, and then the data in the second scene image data set and the first fused data set are input into the first network (i.e., the first network model Net1) in the initial twin network model to obtain the first output result, and at the same time, the data in the second scene image data set and the second fused data set are input into the second network (i.e., the second network model Net2) in the initial twin network model to obtain the second output result. Among them, the number of first scene image data contained in the first fused data set is greater than the number of data in the second scene image data set, and the number of second scene image data contained in the second fused data set is greater than the number of data in the second scene image data set.
[0084] After the above data is input into the two network models of the initial twin network model, the two network models output the first output result and the second output result respectively, and then determine the target twin network model based on the first output result, the second output result and the preset function. Among them, the first output result includes the sub-first output result O_s corresponding to the second scene image data set and the sub-second output result M_s corresponding to the first fusion data set, and the second output result includes the sub-third output result O_t corresponding to the second scene image data set and the sub-fourth output result M_t corresponding to the second fusion data set. Among them, the output shape of O_s and O_t at this time is (B, C), where B represents batch (the number of images input into the model at one time during model training), and C represents the number of distortion categories (i.e., the number of image quality types).
[0085] Based on the first output result, the second output result and the preset function, the target twin network model is determined. First, the sub-fifth output result and the sub-sixth output result need to be determined based on the sub-first output result O_s and the sub-third output result O_t, respectively. Specifically, the image quality type corresponding to the image with the highest score in the sub-first output result O_s and the sub-third output result O_t is selected to obtain the first image quality type P_s and the second image quality type P_t. Here, the maximum (Top1) of the sub-first output result O_s and the sub-third output result O_t in the C dimension is taken, that is, the scores corresponding to all image quality types in the sub-first output result O_s and the sub-third output result O_t are traversed, and the image quality type corresponding to the image with the highest score is found in the sub-first output result O_s and the sub-third output result O_t, respectively. Then the first image quality type P_s, the score S_s corresponding to the first image quality type, the label data Y_s and the threshold T_s are taken as the sub-fifth output result; the second image quality type P_t, the score S_t corresponding to the second image quality type, the label data Y_s and the threshold T_t are taken as the sub-sixth output result.
[0086] After all parameters are determined, the parameters can be iteratively updated through the preset loss function. After the parameters are iteratively updated, the training of the initial twin network model is completed, and the target twin network model is obtained. Figure 7 To illustrate the process of iterative parameter update by presetting the loss function, the details are as follows:
[0087] First, the preset loss function is used to iteratively update the sub-first output result O_s, the sub-second output result M_s, the sub-third output result O_t, the sub-fourth output result M_t, the sub-fifth output result and the sub-sixth output result to obtain an initial fitting result. Specifically, the preset loss function is used to iteratively update the sub-first output result O_s, the sub-second output result M_s, the sub-third output result O_t, the sub-fourth output result M_t, the first image quality type P_s, the score S_s corresponding to the first image quality type, the label data Y_s and the threshold T_s, and the second image quality type P_t, the score S_t corresponding to the second image quality type, the label data Y_s and the threshold T_t to obtain an initial fitting result. Among them, the preset loss function includes the first loss function Loss1, the second loss function Loss2, the third loss function Loss3, the fourth loss function Loss4 and the fifth loss function Loss5, and the initial fitting result includes the first fitting result, the second fitting result, the third fitting result, the fourth fitting result and the fifth fitting result.
[0088] Furthermore, the first fitting result, the second fitting result, the third fitting result, the fourth fitting result and the fifth fitting result are determined as follows:
[0089] Since the target dataset has no artificial labels, the first image quality type P_s and the second image quality type P_t in the following formula are used as pseudo labels.
[0090] (1) Input the sub-second output result M_s, the first image quality type P_s, and the label data Y_s corresponding to the first image quality type into the first loss function Loss1 to obtain a first fitting result. The first fitting result is the value corresponding to Loss1. The calculation formula of Loss1 is as follows:
[0091]
[0092] Among them, B represents batch, that is, the number of images input into the model at one time during model training, i is the i-th image, and α (alpha) is the mixup parameter, which ranges from 0 to 1 and is generally set manually, such as 0.3, 0.5, etc.
[0093] (2) Input the third output result O_t, the first image quality type P_s, and the score S_s and threshold T_s corresponding to the first image quality type into the second loss function Loss2 to obtain a second fitting result. The second fitting result is the value corresponding to Loss2. The calculation formula of Loss2 is as follows:
[0094]
[0095] Among them, B represents batch, that is, the number of images input into the model at one time during model training, and i is the i-th image.
[0096] (3) Input the sub-first output result O_s, the second image quality type P_t, and the score S_t and threshold T_t corresponding to the second image quality type into the third loss function Loss3 to obtain a third fitting result. The third fitting result is the value corresponding to Loss3. The calculation formula of Loss3 is as follows:
[0097]
[0098] Among them, B represents batch, that is, the number of images input into the model at one time during model training, and i is the i-th image.
[0099] (4) Input the fourth output result M_t, the second image quality type P_t, and the label data Y_s corresponding to the second image quality type into the fourth loss function Loss4 to obtain a fourth fitting result. The fourth fitting result is the value corresponding to Loss4. The calculation formula of Loss4 is as follows:
[0100]
[0101] Among them, B represents batch, that is, the number of images input into the model at one time during model training, i is the i-th image, and α (alpha) is the mixup parameter, which ranges from 0 to 1 and is generally set manually, such as 0.3, 0.5, etc.
[0102] (5) Input the sub-first output result O_s and the sub-third output result O_t into the fifth loss function Loss5 to obtain the fifth fitting result. The fifth fitting result is the value corresponding to Loss5. The calculation formula of Loss5 is as follows:
[0103]
[0104] Among them, B represents batch, that is, the number of images input into the model at one time during model training, and i is the i-th image.
[0105] After obtaining the initial fitting result through the above-mentioned preset loss function, it is necessary to use the initial twin network model as the target twin network model when the initial fitting result meets the preset fitting result.
[0106] Step S404: Determine the image quality of the original image based on the multiple sub-images and the target twin network model.
[0107] After the target twin network model is determined through the above steps, multiple sub-images need to be input into the first network in the target twin network model to obtain multiple target results, wherein the multiple sub-images correspond to the multiple target results one by one. Then, the score corresponding to the image quality type is extracted from each of the multiple target results to obtain the score corresponding to the image quality type corresponding to each of the multiple sub-images, and then the score corresponding to the image quality type corresponding to each sub-image is compared with the preset standard score to determine the image quality type corresponding to each sub-image, and finally, based on the image quality type corresponding to each sub-image, the image quality of the original image is determined.
[0108] An embodiment of the present invention provides a method for determining image quality, including: receiving an original image, then preprocessing the original image to obtain multiple sub-images, then using a first scene image data set and a second scene image data set to train an initial twin network model, determine a target twin network model, and finally determine the image quality of the original image based on the multiple sub-images and the target twin network model. The target twin network model of the present invention includes two network models. The two network models enhance the generalization ability of the two networks for image data in different scenes through mutual supervised learning. In addition, the target twin network model can support quality detection of images of different sizes. When the original image is an image of any scene, the target twin network model can effectively output its image quality type, and the determination of image quality will not be affected by the image size problem, thereby improving the accuracy of image quality determination.
[0109] It should be understood that the order of execution of the steps in the above embodiment does not necessarily mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiment of the present invention.
[0110] The following is an embodiment of the device of the present invention. For details not described in detail therein, reference may be made to the corresponding method embodiment described above.
[0111] Figure 8 A schematic diagram of the structure of an image quality determination device provided by an embodiment of the present invention is shown. For the convenience of description, only the part related to the embodiment of the present invention is shown. An image quality determination device includes an image receiving module 81, an image preprocessing module 82, a model training module 83 and an image quality determination module 84, which are specifically as follows:
[0112] An image receiving module 81 is used to receive an original image;
[0113] An image preprocessing module 82 is used to preprocess the original image to obtain multiple sub-images;
[0114] A model training module 83 is used to train the initial twin network model using the first scene image data set and the second scene image data set to determine the target twin network model, wherein the first scene image data set includes image quality data corresponding to the first scene image;
[0115] The image quality determination module 84 is used to determine the image quality of the original image based on multiple sub-images and the target twin network model.
[0116] In a possible implementation, the image preprocessing module 82 includes:
[0117] The format conversion submodule is used to convert the format of the original image to obtain the original image after the format conversion;
[0118] The sliding window processing submodule is used to perform sliding window processing on the original image after format conversion using a preset sliding window to obtain multiple sub-images.
[0119] In a possible implementation, the model training module 83 includes:
[0120] A data fusion submodule, used to fuse the first scene image data set and the second scene image data set to obtain a fused data set;
[0121] a data extraction submodule, configured to extract data from the fused data set to obtain a first fused data set and a second fused data set, wherein the amount of first scene image data included in the first fused data set is greater than the amount of data in the second scene image data set, and the amount of second scene image data included in the second fused data set is greater than the amount of data in the second scene image data set;
[0122] A network processing submodule, used to input the data in the second scene image data set and the first fused data set into the first network in the initial twin network model to obtain a first output result, and simultaneously input the data in the second scene image data set and the second fused data set into the second network in the initial twin network model to obtain a second output result;
[0123] The function processing submodule is used to determine the target twin network model based on the first output result, the second output result and the preset loss function.
[0124] In a possible implementation, the first output result includes a sub-first output result corresponding to the second scene image dataset and a sub-second output result corresponding to the first fused dataset, and the second output result includes a sub-third output result corresponding to the second scene image dataset and a sub-fourth output result corresponding to the second fused dataset;
[0125] The function processing submodules include:
[0126] a calculation unit, configured to determine a sub-fifth output result and a sub-sixth output result based on the sub-first output result and the sub-third output result respectively;
[0127] a parameter updating unit, configured to iteratively update the sub-first output result, the sub-second output result, the sub-third output result, the sub-fourth output result, the sub-fifth output result, and the sub-sixth output result using a preset loss function to obtain an initial fitting result;
[0128] The fitting unit is used to use the initial twin network model as the target twin network model when the initial fitting result meets the preset fitting result.
[0129] In a possible implementation, the computing unit includes:
[0130] An image selection subunit, used to select the image quality type corresponding to the image with the highest score in the sub-first output result and the sub-third output result respectively, to obtain a first image quality type and a second image quality type;
[0131] A first output result determination subunit, configured to use the first image quality type, the score corresponding to the first image quality type, the label data, and the threshold as a fifth sub-output result;
[0132] The second output result determination subunit is used to use the second image quality type, the score corresponding to the second image quality type, the label data and the threshold as the sixth sub output result.
[0133] In a possible implementation, the preset loss function includes a first loss function, a second loss function, a third loss function, a fourth loss function, and a fifth loss function, and the initial fitting result includes a first fitting result, a second fitting result, a third fitting result, a fourth fitting result, and a fifth fitting result;
[0134] The parameter updating unit includes:
[0135] A first parameter updating subunit is used to input the sub-second output result, the first image quality type, and the label data corresponding to the first image quality type into a first loss function to obtain a first fitting result;
[0136] A second parameter updating subunit is used to input the sub-third output result, the first image quality type, and the score and threshold corresponding to the first image quality type into a second loss function to obtain a second fitting result;
[0137] A third parameter updating subunit is used to input the first output result, the second image quality type, and the score and threshold corresponding to the second image quality type into a third loss function to obtain a third fitting result;
[0138] a fourth parameter updating subunit, configured to input the fourth output result, the second image quality type, and the label data corresponding to the second image quality type into a fourth loss function to obtain a fourth fitting result;
[0139] The fifth parameter updating subunit is used to input the sub-first output result and the sub-third output result into the fifth loss function to obtain a fifth fitting result.
[0140] In a possible implementation, the image quality determination module 84 includes:
[0141] A model prediction submodule, used for inputting a plurality of sub-images into a first network in a target twin network model to obtain a plurality of target results, wherein the plurality of sub-images correspond one-to-one to the plurality of target results;
[0142] A score determination submodule, used to extract a score corresponding to an image quality type from each of the multiple target results, and obtain a score corresponding to an image quality type corresponding to each of the multiple sub-images;
[0143] A score comparison submodule, used to compare the score corresponding to the image quality type corresponding to each sub-image with a preset standard score to determine the image quality type corresponding to each sub-image;
[0144] The image quality determination submodule is used to determine the image quality of the original image based on the image quality type corresponding to each sub-image.
[0145] In a possible implementation manner, the original image is an RGB image.
[0146] Fig. 9 is a schematic diagram of a terminal provided by an embodiment of the present invention. Fig. 9 As shown, the terminal 9 of this embodiment includes: a processor 91, a memory 92, and a computer program 93 stored in the memory 92 and executable on the processor 91. When the processor 91 executes the computer program 93, the steps in the above-mentioned methods for determining image quality are implemented, for example Figure 4 Alternatively, when the processor 91 executes the computer program 93, the functions of each module / unit in the above-mentioned image quality determination device embodiments are implemented, for example Figure 8 Functionality of modules / units 81 to 84 shown.
[0147] The present invention also provides a readable storage medium, in which a computer program is stored. When the computer program is executed by a processor, it is used to implement the image quality determination method provided by the various embodiments described above.
[0148] Among them, the readable storage medium can be a computer storage medium or a communication medium. The communication medium includes any medium that facilitates the transmission of a computer program from one place to another. The computer storage medium can be any available medium that can be accessed by a general-purpose or special-purpose computer. For example, a readable storage medium is coupled to a processor so that the processor can read information from the readable storage medium and write information to the readable storage medium. Of course, the readable storage medium can also be a component of the processor. The processor and the readable storage medium can be located in an application-specific integrated circuit (Application Specific Integrated Circuits, abbreviated as: ASIC). In addition, the ASIC can be located in a user device. Of course, the processor and the readable storage medium can also exist in a communication device as discrete components. The readable storage medium can be a read-only memory (ROM), a random access memory (RAM), a CD-ROM, a magnetic tape, a floppy disk, and an optical data storage device, etc.
[0149] The present invention also provides a program product, which includes an execution instruction, which is stored in a readable storage medium. At least one processor of a device can read the execution instruction from the readable storage medium, and at least one processor executes the execution instruction so that the device implements the image quality determination method provided by the various embodiments described above.
[0150] In the embodiments of the above-mentioned devices, it should be understood that the processor may be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), etc. The general-purpose processor may be a microprocessor or any conventional processor. The steps of the method disclosed in the present invention may be directly implemented as being executed by a hardware processor, or may be executed by a combination of hardware and software modules in the processor.
[0151] The above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit the same. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that the technical solutions described in the aforementioned embodiments may still be modified, or some of the technical features may be replaced by equivalents. Such modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present invention, and should be included in the protection scope of the present invention.
Claims
1. A method for determining image quality, characterized in that: include: receiving the original image; Preprocessing the original image to obtain multiple sub-images; Using a first scene image data set and a second scene image data set to train an initial twin network model to determine a target twin network model, wherein the first scene image data set includes image quality data corresponding to the first scene image; Determining the image quality of the original image based on the multiple sub-images and the target twin network model; The preprocessing of the original image to obtain a plurality of sub-images includes: Converting the format of the original image to obtain an original image after the format conversion; Using a preset sliding window to perform sliding window processing on the original image after the format conversion to obtain a plurality of sub-images; The method of training the initial twin network model using the first scene image data set and the second scene image data set to determine the target twin network model includes: Fusing the first scene image dataset and the second scene image dataset to obtain a fused dataset; Extracting data from the fused data set to obtain a first fused data set and a second fused data set, wherein the amount of first scene image data included in the first fused data set is greater than the amount of second scene image data, and the amount of second scene image data included in the second fused data set is greater than the amount of first scene image data; Inputting the data in the second scene image data set and the first fused data set into the first network in the initial twin network model to obtain a first output result, and simultaneously inputting the data in the second scene image data set and the second fused data set into the second network in the initial twin network model to obtain a second output result; Based on the first output result, the second output result and a preset loss function, a target twin network model is determined.
2. The method for determining image quality according to claim 1, characterized in that: The first output result includes a sub-first output result corresponding to the second scene image dataset and a sub-second output result corresponding to the first fused dataset, and the second output result includes a sub-third output result corresponding to the second scene image dataset and a sub-fourth output result corresponding to the second fused dataset; The determining a target twin network model based on the first output result, the second output result and a preset function includes: Determine a fifth sub-output result and a sixth sub-output result based on the first sub-output result and the third sub-output result respectively; Iteratively updating the sub-first output result, the sub-second output result, the sub-third output result, the sub-fourth output result, the sub-fifth output result, and the sub-sixth output result by using the preset loss function to obtain an initial fitting result; When the initial fitting result meets the preset fitting result, the initial twin network model is used as the target twin network model.
3. The method for determining image quality according to claim 2, characterized in that: The determining the sub-fifth output result and the sub-sixth output result based on the sub-first output result and the sub-third output result respectively comprises: Selecting image quality types corresponding to images with the highest scores in the sub-first output result and the sub-third output result respectively, to obtain a first image quality type and a second image quality type; taking the first image quality type, the score corresponding to the first image quality type, the label data, and the threshold as the fifth sub-output result; The second image quality type, the score corresponding to the second image quality type, the label data and the threshold are used as the sub-sixth output result.
4. The method for determining image quality according to claim 3, characterized in that: The preset loss function includes a first loss function, a second loss function, a third loss function, a fourth loss function and a fifth loss function, and the initial fitting result includes a first fitting result, a second fitting result, a third fitting result, a fourth fitting result and a fifth fitting result; The using the preset loss function to iteratively update the sub-first output result, the sub-second output result, the sub-third output result, the sub-fourth output result, the sub-fifth output result and the sub-sixth output result to obtain an initial fitting result includes: Inputting the sub-second output result, the first image quality type, and label data corresponding to the first image quality type into the first loss function to obtain the first fitting result; Inputting the sub-third output result, the first image quality type, and the score and threshold corresponding to the first image quality type into the second loss function to obtain the second fitting result; Inputting the sub-first output result, the second image quality type, and the score and threshold corresponding to the second image quality type into the third loss function to obtain the third fitting result; Inputting the fourth sub-output result, the second image quality type, and label data corresponding to the second image quality type into the fourth loss function to obtain the fourth fitting result; The sub-first output result and the sub-third output result are input into the fifth loss function to obtain the fifth fitting result.
5. The method for determining image quality according to claim 4, characterized in that: The determining the image quality of the original image based on the multiple sub-images and the target twin network model includes: Inputting the plurality of sub-images into a first network in the target twin network model to obtain a plurality of target results, wherein the plurality of sub-images correspond one-to-one to the plurality of target results; Extracting a score corresponding to an image quality type from each of the multiple target results to obtain a score corresponding to an image quality type corresponding to each of the multiple sub-images; Compare the score corresponding to the image quality type corresponding to each sub-image with a preset standard score to determine the image quality type corresponding to each sub-image; The image quality of the original image is determined based on the image quality type corresponding to each sub-image.
6. The method for determining image quality according to any one of claims 1 to 5, characterized in that: The original image is an RGB image.
7. A device for determining image quality, characterized in that: include: An image receiving module, used for receiving an original image; An image preprocessing module, used for preprocessing the original image to obtain multiple sub-images; A model training module, used to train an initial twin network model using a first scene image data set and a second scene image data set to determine a target twin network model, wherein the first scene image data set includes image quality data corresponding to the first scene image; An image quality determination module, configured to determine the image quality of the original image based on the multiple sub-images and the target twin network model; The preprocessing of the original image to obtain a plurality of sub-images includes: Converting the format of the original image to obtain an original image after the format conversion; Using a preset sliding window to perform sliding window processing on the original image after the format conversion to obtain a plurality of sub-images; The method of training the initial twin network model using the first scene image data set and the second scene image data set to determine the target twin network model includes: Fusing the first scene image dataset and the second scene image dataset to obtain a fused dataset; Extracting data from the fused data set to obtain a first fused data set and a second fused data set, wherein the amount of first scene image data included in the first fused data set is greater than the amount of second scene image data, and the amount of second scene image data included in the second fused data set is greater than the amount of first scene image data; Inputting the data in the second scene image data set and the first fused data set into the first network in the initial twin network model to obtain a first output result, and simultaneously inputting the data in the second scene image data set and the second fused data set into the second network in the initial twin network model to obtain a second output result; Based on the first output result, the second output result and a preset loss function, a target twin network model is determined.
8. A terminal comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, the steps of the method for determining image quality according to any one of claims 1 to 6 are implemented.
9. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the computer program implements the steps of the method for determining image quality according to any one of claims 1 to 6.
Citation Information
Patent Citations
A contrast learning image quality evaluation method based on a twin network
CN109727246A
Mobile phone screen defect detection method, device and system, computer equipment and medium
CN111612763A