Image exposure quantification method and electronic device

By using an image-text comparison model to obtain the semantic similarity between an image and preset prompts, the problem of low accuracy in traditional image exposure determination methods is solved, thereby improving the accuracy and efficiency of image exposure quantification.

CN119559453BActive Publication Date: 2025-10-24HONOR DEVICE CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510130356.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-02-05
Publication Date
2025-10-24
Estimated Expiration
2045-02-05

AI Technical Summary

Technical Problem

Traditional image exposure determination methods suffer from low accuracy, especially since manual determination involves subjective ambiguity and a complex process.

Method used

By employing a pre-trained image-text comparison model, the exposure status label of an image is determined by obtaining the semantic similarity value between the input image and the preset prompt words, thus avoiding the subjective ambiguity and complex process of manual evaluation.

Benefits of technology

It improves the accuracy and efficiency of image exposure quantization results, reduces computing power requirements, and simplifies the quantization process.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119559453B_ABST
    Figure CN119559453B_ABST
Patent Text Reader

Abstract

The application discloses an image exposure quantification method and an electronic device, and relates to the technical field of image processing, and comprises: acquiring an input image. Wherein, the input image comprises one or more images to be subjected to exposure quantification. The input image is input into a picture-text contrast model to obtain a semantic similarity value of the input image and a preset prompt word. Wherein, the preset prompt word comprises one or more description fields representing exposure states. Based on the semantic similarity value of the input image corresponding to the preset prompt word, an exposure quantification label of the input image is determined. Wherein, the exposure quantification label is used to represent that the input image is an underexposed image, a normally exposed image or an overexposed image. In the scheme, the picture-text contrast model can learn the semantic relationship between the preset prompt word (text) and the image, so as to obtain the semantic similarity value of the image and the preset prompt word. The exposure state label of the image based on the semantic similarity value is objective and accurate.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] Embodiments of the present application relate to the technical field of image processing, and in particular to an image exposure quantification method and an electronic device. BACKGROUND

[0002] Image exposure refers to the process in which a light-sensitive element (such as a camera sensor or film) receives light and converts it into an image signal during image acquisition. Image exposure involves multiple factors such as the intensity of light, exposure time, and the sensitivity of the light-sensitive element, which have a decisive influence on the brightness, contrast, and detail performance of the image. Overexposure of an image will result in the overall image being too bright, with loss of detail in the highlight part; underexposure of an image will result in the overall image being too dark, with loss of detail in the shadow part.

[0003] In image processing scenarios or some shooting scenarios, it is necessary to make exposure judgments such as underexposure, normal exposure, or overexposure of an image. The exposure determination method provided by the traditional technology, such as manual exposure determination, has the problem of low determination accuracy. SUMMARY

[0004] Embodiments of the present application provide an image exposure quantification method and an electronic device, which can obtain a similarity value between an input image and a prompt word through a trained image-text contrast model. The prompt word includes a field for representing different exposure states. Based on the similarity value, the exposure state label of the input image can be determined, avoiding the exposure determination inaccuracy caused by the subjective ambiguity of manual exposure determination, and the complex process of exposure determination by finding a reference and the inaccuracy of exposure determination, so that the exposure quantification result of the input image is relatively accurate.

[0005] To achieve the above-mentioned purpose, embodiments of the present application adopt the following technical solutions.

[0006] In a first aspect, an image exposure quantification method is provided, comprising:

[0007] An input image is obtained. The input image includes one or more images to be subjected to exposure quantification.

[0008] The input image is input into an image-text contrast model to obtain a semantic similarity value between the input image and a preset prompt word. The preset prompt word includes one or more description fields representing exposure states.

[0009] The model is obtained by iterative training based on the prediction of the exposure quantification label (training label) of a sample image by the sample image and a training prompt word, and the loss of the training label and the reference label (actual exposure state) of the sample image. The model can output the semantic similarity between the image features and the text (prompt word) features, thereby achieving the purpose of understanding the exposure state of the image.

[0010] determine the exposure quantification label of the input image based on the semantic similarity value between the input image and the preset prompt word. The exposure quantification label is used to represent that the input image is an underexposed image, a normally exposed image or an overexposed image.

[0011] The electronic device can be used as an execution subject of the image exposure quantification method.

[0012] In the present application, the semantic similarity value between the input image and the preset prompt word is obtained by using the trained image-text comparison model. The preset prompt word includes fields describing different exposure states. The exposure state label of the input image is determined based on the semantic similarity value. The image-text comparison model can learn the semantic relationship between the preset prompt word (text) and the image, so as to obtain the semantic similarity value between the image and the preset prompt word. The exposure state label of the image obtained based on the semantic similarity value is objective and accurate. The subjective ambiguity of manual evaluation is avoided, and the entire process does not need to find reference benchmark images or develop complex multiple quantification strategies. The exposure quantification accuracy is ensured, and the computing power is reduced.

[0013] In a possible implementation manner of the first aspect, the input image includes one image to be quantified, and the preset prompt word includes multiple prompt words describing different exposure states.

[0014] The electronic device inputs the input image into the image-text comparison model to obtain the semantic similarity value between the input image and the preset prompt word, including:

[0015] The electronic device inputs the input image into the image-text comparison model to obtain the first similarity value corresponding to each preset prompt word of the input image, and obtains multiple first similarity values.

[0016] Then, the electronic device determines the exposure quantification label of the input image based on the semantic similarity value, including:

[0017] The electronic device determines the exposure quantification label of the input image based on the exposure state corresponding to the preset prompt word with the maximum first similarity value.

[0018] In the present application, the semantic similarity value between the input image and the preset prompt word is obtained by the trained image-text contrast model, wherein the preset prompt word includes a field describing different exposure states. Based on the semantic similarity value, the exposure state label of the input image can be determined. Wherein the image-text contrast model can learn the semantic relationship between the preset prompt word (text) and the image, so as to obtain the semantic similarity value of the image and the preset prompt word. The greater the first similarity value is, the higher the semantic similarity between the image and the preset prompt word is. The exposure state corresponding to the preset prompt word with the maximum first similarity value is determined as the exposure quantization label of the input image, which is objective and accurate. Avoiding the subjective ambiguity of manual evaluation, the entire process does not need to find reference benchmark images, nor does it need to develop complex multiple quantization strategies, while ensuring the accuracy of exposure quantization, it also reduces the computing power.

[0019] In another possible implementation of the first aspect, the input image includes an image to be exposed quantization, and the preset prompt word includes a plurality of prompt words describing different exposure states.

[0020] The method further includes:

[0021] The electronic device performs significant model inference on the input image to obtain a subject image corresponding to the input image.

[0022] The electronic device inputs the subject image into the image-text contrast model to obtain a second similarity value corresponding to each preset prompt word of the subject image.

[0023] Then, the electronic device inputs the input image into the image-text contrast model to obtain the semantic similarity value between the input image and the preset prompt word, including:

[0024] The electronic device inputs the input image into the image-text contrast model to obtain a first similarity value corresponding to each preset prompt word of the input image, thereby obtaining a plurality of first similarity values.

[0025] Then, the electronic device determines the exposure quantization label of the input image based on the semantic similarity value, including:

[0026] Based on the first similarity value and the second similarity value, a third similarity value of the input image is determined, and the electronic device determines the exposure quantization label of the input image based on the third similarity value.

[0027] In the present application, the trained image-text contrast model can realize semantic understanding of the relevance of the global image of the input image to the preset prompt word, and semantic understanding of the relevance of the subject image corresponding to the input image to the preset prompt word. For input images containing some special scenes, performing semantic understanding of the relevance of the subject image to the preset prompt word can better determine the actual exposure state of the input image, so that the exposure quantization label of the final obtained input image is more accurate.

[0028] In a possible implementation form of the first aspect, the first similarity value comprises a semantic similarity value of the input image corresponding to the preset prompt word, and the second similarity value comprises a semantic similarity value of the subject image corresponding to the preset prompt word.

[0029] Based on the first similarity value and the second similarity value, a third similarity value of the input image is determined, comprising:

[0030] For each preset prompt word, the first similarity value and the second similarity value are weighted and summed to obtain a third similarity value corresponding to each preset prompt word, and a plurality of third similarity values are obtained.

[0031] Then, the electronic device determines the exposure quantization label of the input image based on the third similarity value, comprising:

[0032] The electronic device determines the exposure quantization label of the input image based on the exposure state corresponding to the preset prompt word with the maximum third similarity value.

[0033] The weights corresponding to the first similarity value and the second similarity value in the weighted calculation can be determined according to actual conditions.

[0034] In the present application, the first similarity value and the second similarity value can be weighted and summed to obtain a third similarity value corresponding to each preset prompt word, and a plurality of third similarity values are obtained. The exposure quantization label of the input image is determined based on the exposure state corresponding to the preset prompt word with the maximum third similarity value. The third similarity value combines the semantic understanding result (similarity value) of the input image (global image) and the semantic understanding result (similarity value) of the subject image. The third similarity value can more accurately represent the exposure state of the input image, and the exposure quantization label of the input image obtained based on the third similarity value is more accurate.

[0035] In a possible implementation form of the first aspect, the first similarity value comprises a semantic similarity value of the input image corresponding to the preset prompt word, and the second similarity value comprises a semantic similarity value of the subject image corresponding to the preset prompt word.

[0036] Based on the first similarity value and the second similarity value, a third similarity value of the input image is determined, comprising:

[0037] The electronic device obtains a first similarity value with the largest value and a second similarity value with the largest value.

[0038] In a case where the preset prompt corresponding to the first similarity value with the largest value is consistent with the preset prompt corresponding to the second similarity value with the largest value, the electronic device performs weighted summation on the first similarity value with the largest value and the second similarity value with the largest value to obtain a third similarity value.

[0039] Then, the electronic device determines the exposure quantization label of the input image based on the third similarity value, including:

[0040] The electronic device determines the exposure quantization label of the input image based on the preset prompt corresponding to the third similarity value.

[0041] Generally, the similarity trend of the global image to the preset prompt is consistent with the similarity trend of the local image to the preset prompt.

[0042] In the present application, in a case where the preset prompt corresponding to the first similarity value with the largest value is consistent with the preset prompt corresponding to the second similarity value with the largest value, the electronic device performs weighted summation on the first similarity value with the largest value and the second similarity value with the largest value to obtain a third similarity value. The third similarity value combines the semantic understanding result (similarity value) of the input image (global image) and the semantic understanding result (similarity value) of the subject image, and can more accurately represent the exposure state of the subject in the input image. The exposure quantization label of the input image obtained based on the third similarity value is more accurate.

[0043] In another possible implementation of the first aspect, the method further includes:

[0044] If the preset prompt corresponding to the first similarity value with the largest value is inconsistent with the preset prompt corresponding to the second similarity value with the largest value, the electronic device takes the second similarity value with the largest value as the third similarity value.

[0045] In some other cases, there can also be a case where the similarity trend of the global image to the preset prompt is inconsistent with the similarity trend of the local image to the preset prompt.

[0046] In the present application, in a case where the preset prompt corresponding to the first similarity value with the largest value is inconsistent with the preset prompt corresponding to the second similarity value with the largest value, the electronic device takes the second similarity value with the largest value as the third similarity value. The third similarity value takes the semantic understanding result (similarity value) of the subject image as the criterion, and focuses more on the exposure quantization of the subject of the image. The third similarity value can more accurately represent the exposure state of the input image, and the exposure quantization label of the input image obtained based on the third similarity value is more accurate.

[0047] In a possible implementation of the first aspect, the preset prompt words include a first prompt word, a second prompt word, and a third prompt word, the first prompt word is a prompt word representing underexposure, the second prompt word is a prompt word representing normal exposure, and the third prompt word is a prompt word representing overexposure.

[0048] In the present application, the preset prompt words can include prompt words representing different image exposure states, and the text feature learning of the image-text contrast model can be performed on the prompt words representing different image exposure states, so as to achieve a more comprehensive understanding of the exposure state. When the image-text contrast model is used for semantic understanding of the image, the exposure state of the image can be more comprehensively matched, so that the similarity of the output image to the preset prompt word is more accurate, and the exposure quantization label of the image is more in line with the actual exposure state of the image, and is more accurate.

[0049] In a possible implementation of the first aspect, the input image includes a plurality of images to be quantized in exposure, and the preset prompt words include a target prompt word of a target exposure state.

[0050] The electronic device inputs the input image into the image-text contrast model to obtain a semantic similarity value of the input image to the preset prompt word, including:

[0051] The electronic device inputs each input image into the image-text contrast model to obtain a fourth similarity value of each input image corresponding to the target prompt word.

[0052] Then, the electronic device determines an exposure quantization label of the input image based on the semantic similarity value, including:

[0053] The electronic device determines an exposure quantization label of each input image based on the fourth similarity value of each input image and a preset threshold strategy.

[0054] In the present application, the semantic understanding of the image and the target prompt word by the trained image-text contrast model can also be used to realize the query of the target image based on the target prompt word. For example, based on the similarity value of each input image to the target prompt word and the preset threshold strategy, the exposure quantization label of each input image can be determined, and the image closer to the target prompt word can be determined as the target image from the plurality of input images. The exposure quantization of the input image is obtained by the image-text contrast model, and the quantization result is objective and accurate, which can meet the image exposure quantization demand in various scenes.

[0055] In a possible implementation of the first aspect, the preset threshold strategy is related to the target prompt word.

[0056] The preset threshold strategy includes:

[0057] When the target prompt word is a prompt word representing normal exposure, the exposure quantization label of the input image with the fourth similarity value greater than or equal to the first threshold value is a normal exposure image. The exposure quantization label of the input image with the fourth similarity value less than the first threshold value is a non-normal exposure image.

[0058] When the target prompt word is a prompt word representing underexposure, the exposure quantization label of the input image with the fourth similarity value greater than or equal to the second threshold value is an underexposure image. The exposure quantization label of the input image with the fourth similarity value less than the second threshold value and greater than or equal to the third threshold value is a normal exposure image. The exposure quantization label of the input image with the fourth similarity value less than the third threshold value is an overexposure image.

[0059] When the target prompt word is a prompt word representing overexposure, the exposure quantization label of the input image with the fourth similarity value greater than or equal to the fourth threshold value is an overexposure image. The exposure quantization label of the input image with the fourth similarity value less than the fourth threshold value and greater than or equal to the fifth threshold value is a normal exposure image. The exposure quantization label of the input image with the fourth similarity value less than the fifth threshold value is an underexposure image.

[0060] In the present application, the semantic understanding of the correlation between the image and the target prompt word of the trained image-text contrast model can also be used to realize the query of the target image based on the target prompt word. Based on the similarity values of the input images and the target prompt word and the threshold value strategies corresponding to different target prompt words, the image closer to the target prompt word can be determined as the target image from the plurality of input images by the method provided in the above embodiments. The exposure quantization of the input image is obtained by the image-text contrast model, and the quantization result is objective and accurate, which can meet the image exposure quantization demand in various scenarios.

[0061] In a possible implementation manner of the first aspect, the input image includes a plurality of images to be subjected to exposure quantization, and the preset prompt word includes a target prompt word of a target exposure state.

[0062] The method further includes:

[0063] The electronic device performs saliency model inference on each input image to obtain a subject image corresponding to each input image.

[0064] The electronic device inputs each subject image into the image-text contrast model to obtain a fifth similarity value of each subject image corresponding to the target prompt word.

[0065] Then, the electronic device inputs the input image into the image-text contrast model to obtain a semantic similarity value of the input image and the preset prompt word, including:

[0066] The electronic device inputs each input image into the image-text contrast model to obtain a fourth similarity value corresponding to the target prompt word for each input image.

[0067] Then, the electronic device determines an exposure quantization label of the input image based on the semantic similarity value, including:

[0068] The electronic device determines a sixth similarity value of the input image based on the fourth similarity value corresponding to each input image and the fifth similarity value of the subject image corresponding to the input image.

[0069] The electronic device determines an exposure quantization label of each input image based on the sixth similarity value of each input image and a preset exposure threshold.

[0070] In the present application, the semantic understanding of the correlation between the image and the target prompt word of the trained image-text contrast model can also be used to realize the query of the target image based on the target prompt word. For example, when the target image is a normally exposed image, the sixth similarity value can be determined based on the similarity value of each input image and the target prompt word and the similarity value of each subject image and the target prompt word, and the exposure quantization label of each input image can be determined based on the sixth similarity value and a preset threshold strategy, and the image closer to the target prompt word is determined as the target image from the multiple input images. The exposure quantization of the input image is obtained through the image-text contrast model, and the quantization result is objective and accurate, which can meet the image exposure quantization demand in various scenes.

[0071] In a possible implementation form of the first aspect, the electronic device determines a sixth similarity value of the input image based on the fourth similarity value corresponding to each input image and the fifth similarity value of the subject image corresponding to the input image, including:

[0072] The electronic device takes the average value of the fourth similarity value corresponding to each input image and the fifth similarity value of the subject image corresponding to the input image as the sixth similarity value of the input image.

[0073] Alternatively,

[0074] The electronic device takes the weighted sum value of the fourth similarity value corresponding to each input image and the fifth similarity value of the subject image corresponding to the input image as the sixth similarity value of the input image.

[0075] In the present application, the sixth similarity value combines the semantic understanding result (similarity value) of the input image (global image) and the semantic understanding result (similarity value) of the subject image, which can more accurately represent the exposure state of the subject in the input image, and the exposure quantization label of the input image based on the sixth similarity value is more accurate, and the target image corresponding to the target prompt word obtained is also more accurate.

[0076] In a possible implementation of the first aspect, the preset threshold strategy is associated with a target prompt word.

[0077] The preset threshold strategy comprises:

[0078] When the target prompt word represents normal exposure, the exposure quantization label of the input image with the sixth similarity value greater than or equal to the first threshold is a normal exposure image. The exposure quantization label of the input image with the sixth similarity value less than the first threshold is a non-normal exposure image.

[0079] When the target prompt word represents underexposure, the exposure quantization label of the input image with the sixth similarity value greater than or equal to the second threshold is an underexposure image. The exposure quantization label of the input image with the sixth similarity value less than the second threshold and greater than or equal to the third threshold is a normal exposure image. The exposure quantization label of the input image with the sixth similarity value less than the third threshold is an overexposure image.

[0080] When the target prompt word represents overexposure, the exposure quantization label of the input image with the sixth similarity value greater than or equal to the fourth threshold is an overexposure image. The exposure quantization label of the input image with the sixth similarity value less than the fourth threshold and greater than or equal to the fifth threshold is a normal exposure image. The exposure quantization label of the input image with the sixth similarity value less than the fifth threshold is an underexposure image.

[0081] In the present application, the image-target prompt word correlation semantic understanding of the trained image-text contrast model can also be used to realize target image query based on the target prompt word. Through the method provided in the above embodiments, the image closer to the target prompt word can be determined as the target image from the multiple input images based on the similarity values of the input images and the target prompt word, the similarity values of the subject images and the target prompt word, and the threshold strategies corresponding to different target prompt words. The exposure quantization of the input image is obtained through the image-text contrast model, and the quantization result is objective and accurate, which can meet the image exposure quantization demand in various scenarios.

[0082] In a second aspect, the present application provides a model training method applied to an electronic device, which further comprises:

[0083] The electronic device obtains a plurality of sample images.

[0084] The electronic device inputs the plurality of sample images into an initial image-text contrast model to obtain a first training similarity value of each sample image and a plurality of first training prompt words. The plurality of first training prompt words comprise prompt words representing different exposure states.

[0085] The electronic device determines a training label of the sample image based on an exposure state corresponding to the first training prompt word with the maximum first training similarity value.

[0086] The electronic device calculates a loss based on the training label of the sample image and a reference label of the sample image; the reference label is an actual exposure state of the sample image, and the reference label is used to represent that the sample image is an underexposed image, a normally exposed image, or an overexposed image.

[0087] The electronic device iteratively trains the initial image-text contrast model based on the loss until the initial image-text contrast model meets a training condition, to obtain the image-text contrast model.

[0088] In this application, the electronic device can calculate the loss (such as the correlation loss) between the training label of the sample image and the reference label corresponding to the sample image, iteratively train the image-text contrast model based on the loss, adjust the model parameters and the training prompt words, until the training condition is met, so as to obtain the trained image-text contrast model. The training condition can be that the loss gradually approaches 0, or the number of iterations meets the number threshold. Model training can make the image-text contrast model more accurate in understanding the exposure of the image, and make the similarity between the output sample image and the training prompt word more accurate, so that the exposure state label of the image obtained based on the trained image-text contrast model is more accurate.

[0089] The image exposure quantification method provided in this embodiment does not need to develop a complex quantification strategy, and does not need to fine-tune the model based on the reference image, but the trained image-text contrast model can still guarantee the accuracy of the understanding of the exposure state of the image. That is, the image exposure quantification method provided in this embodiment only needs to adjust the model parameters (such as model weights) and the training prompt words in the model training process while ensuring the accuracy of the output results of the image-text contrast model, without fine-tuning the entire model. The model training is simpler and more efficient.

[0090] In a possible implementation form of the second aspect, the method further includes:

[0091] The electronic device performs saliency model inference on the plurality of sample images to obtain a sample subject image corresponding to each sample image.

[0092] The electronic device inputs the sample subject image into the initial image-text contrast model to obtain a second training similarity value between each sample subject image and a plurality of second training prompt words. The plurality of second training prompt words include prompt words representing different exposure states.

[0093] Then, the training label of the sample image is determined based on the exposure state corresponding to the first training prompt word with the maximum first training similarity value, including:

[0094] For each sample image, the electronic device obtains a first training similarity value with the maximum value and a second training similarity value with the maximum value of the sample subject image corresponding to the sample image.

[0095] In a case where the first training prompt word corresponding to the first training similarity value with the maximum value of the sample image and the second training prompt word corresponding to the second training similarity value with the maximum value are consistent in representing the exposure state, the electronic device determines a third training similarity value based on the first training similarity value with the maximum value and the second training similarity value with the maximum value.

[0096] The electronic device determines the training label of the sample image based on the first training prompt word or the second training prompt word corresponding to the third training similarity value.

[0097] In the present application, compared with the conventional image exposure quantification method based on the quality evaluation model, there is no need to formulate a complex quantification strategy, and there is no need to fine-tune the model based on the reference image, but the trained image-text contrast model can still guarantee the accuracy of the understanding of the exposure state of the image. That is, the image exposure quantification method provided in the embodiment only needs to adjust the model parameters (such as model weights) and training prompt words, and does not need to additionally obtain image data of different exposure states for fine-tuning of the model. While ensuring the accuracy of the output result of the image-text contrast model, the model training process is simpler and the model training efficiency is higher. Moreover, in the embodiment, the relevance learning of the local image and the training prompt word is introduced, so that the image-text contrast model can obtain more accurate exposure quantification results for the entire image by exposure quantification of the main body for some images containing special scenes, such as images with large-area light-colored background or large-area dark-colored background, so that the image exposure quantification is more accurate.

[0098] In a further possible implementation form of the second aspect, determining the training label of the sample image based on the exposure state corresponding to the first training prompt word with the maximum first training similarity value further comprises:

[0099] If the preset prompt word corresponding to the maximum first similarity value is inconsistent with the preset prompt word corresponding to the maximum second similarity value, the maximum second similarity value is taken as the third training similarity value.

[0100] The electronic device determines the training label of the sample image based on the first training prompt word or the second training prompt word corresponding to the third training similarity value.

[0101] In the present application, only the model parameters (such as model weights) and training prompt words need to be adjusted, and no additional image data of different exposure states needs to be obtained for model fine-tuning. In this way, the model training process is simpler, and the model training efficiency is higher while ensuring the accuracy of the output results of the image-text contrast model. Moreover, in the present embodiment, the correlation learning of the local image and the training prompt word is introduced, so that the image-text contrast model can obtain more accurate exposure quantization results for the entire image by exposure quantization of the main body for some images containing special scenes, such as images with a large area of light background or a large area of dark background, so that the image exposure quantization is more accurate.

[0102] In a third aspect, an electronic device is provided, which includes a memory, a processor, and a computer program stored in the memory, and the processor executes the computer program to implement the steps of the method of any one of the first aspect.

[0103] In a fourth aspect, a computer readable storage medium is provided, which stores instructions, and the computer program / instructions are executed by a processor to implement the steps of the method of any one of the first aspect.

[0104] In a fifth aspect, a computer program product containing instructions is provided, which includes a computer program / instructions, and the computer program / instructions are executed by a processor to implement the steps of the method of any one of the first aspect.

[0105] In a sixth aspect, an embodiment of the present application provides a chip, which includes a processor configured to invoke a computer program in a memory to execute the method of any one of the first aspect.

[0106] It can be understood that the electronic device of the third aspect, the computer readable storage medium of the fourth aspect, the computer program product of the fifth aspect, and the chip of the sixth aspect can achieve the beneficial effects of the first aspect and any possible design of the first aspect, which will not be described here. BRIEF DESCRIPTION OF DRAWINGS

[0107] Figure 1 is a comparison diagram of a normally exposed image and an underexposed image;

[0108] Figure 2 is a comparison diagram of a normally exposed image and an overexposed image;

[0109] Figure 3 is a diagram of a properly exposed image selected by different observers under a subjective evaluation method;

[0110] Figure 4 is a before-and-after comparison diagram of a corrected exposure of a snow scene;

[0111] Figure 5 A schematic diagram of a training path of a common image exposure quantification quality evaluation model;

[0112] Figure 6 A schematic diagram of the structure of an electronic device provided in an embodiment of the present application;

[0113] Figure 7 A schematic diagram of a group of images in different exposure states provided in an embodiment of the present application;

[0114] Figure 8 A schematic diagram of a fold line of similarity values between different prompt words and images provided in an embodiment of the present application;

[0115] Figure 9 A schematic diagram of a training path of a text-image comparison model provided in an embodiment of the present application;

[0116] Figure 10 A schematic diagram of a flow of an image exposure quantification method provided in an embodiment of the present application;

[0117] Figure 11 A schematic diagram of a training path of another text-image comparison model provided in an embodiment of the present application;

[0118] Figure 12 A schematic diagram of a flow of another image exposure quantification method provided in an embodiment of the present application;

[0119] Figure 13 A schematic diagram of a flow of another image exposure quantification method provided in an embodiment of the present application;

[0120] Figure 14 A comparison schematic diagram of exposure quantification evaluation provided in an embodiment of the present application;

[0121] Figure 15 A schematic diagram of the structure of another electronic device provided in an embodiment of the present application;

[0122] Figure 16 A schematic diagram of the structure of a chip system provided in an embodiment of the present application. DETAILED DESCRIPTION

[0123] In the description of the embodiments of the application, the terms used in the following embodiments are only for the purpose of describing the specific embodiments and are not intended to be limiting to the application. As used in the specification and the appended claims of the application, the singular forms “a,” “an” and “the” are intended to include the plural forms as well, unless the context clearly indicates otherwise. It will be further understood that the terms “comprises” and / or “comprising,” when used in this specification, specify the presence of stated features, integers, steps, operations, elements, and / or components, but do not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof. The term “and / or” used in the context of the application refers to three relationships: A and / or B, A or B, and A and B. For example, A and / or B can mean A alone, A and B together, and B alone, where A and B can be singular or plural.

[0124] Reference in the specification to “one embodiment” or “some embodiments” means that a particular feature, structure, or characteristic described in connection with the embodiment is included in at least one embodiment of the application. The appearances of the phrase “in one embodiment” or “in some embodiments” in various places in the specification are not necessarily all referring to the same embodiment, although it can. The terms “comprise,” “comprising,” “include,” “including,” and “has” or “having” (and variants thereof) are intended to be open-ended terms that specifically permit the presence of one or more other features, integers, steps, operations, elements, and / or components not expressly named. The term “connected” can include both direct and indirect connections. The terms “first,” “second,” and “third” are used merely as labels, and are not intended to signify relative importance or a quantity of the indicated number.

[0125] In the embodiments of the application, the words “exemplary” and “for example” are used to mean serving as an example, instance, or illustration. Any embodiment or design described herein as “exemplary” or “for example” is not necessarily to be construed as preferred or advantageous over other embodiments or designs. Rather, use of the words “exemplary” and “for example” is intended to present concepts in a concrete manner. In the embodiments of the application, the terms “first,” “second,” and “third” are used merely as labels, and are not intended to signify relative importance or a quantity of the indicated number.

[0126] Image exposure refers to the process in which a light-sensitive element, such as a camera sensor or film, receives light and converts it into an image signal during image acquisition. Image exposure involves factors such as light intensity, exposure time, and the sensitivity of the light-sensitive element, which have a decisive influence on the brightness, contrast, and detail performance of the image. Underexposure and overexposure are two common situations of improper image exposure, which can significantly affect the quality and visual effect of the image.

[0127] Specifically, the overall brightness of the image is moderate, neither too bright nor too dark. The details of both the bright part and the dark part can be well presented, such as the facial expression of the person being photographed, the clear skin texture, the details of the wrinkles of the clothes, the gloss of the accessories, etc. In a landscape photo, the details such as the cloud layers in the sky, the contours of the distant mountains, and the leaf veins of the nearby flowers and plants can be preserved, making the image look rich and complete.

[0128] Underexposure refers to insufficient exposure of an image, resulting in an overall dark image and loss of details in the shadow part. In the case of underexposure, the dark part of the image is too deep, and the texture and level cannot be distinguished, while the bright part is relatively normal, but the overall brightness and contrast of the image are low.

[0129] Exemplarily, Figure 1 A comparison diagram of a normally exposed image and an underexposed image is given.

[0130] Figure 1 A comparison diagram of a normally exposed image and an underexposed image of a cake placed on a white table (light background 1) is given. Among them, Figure 1 (a) of FIG. 1 shows a normally exposed image, Figure 1 (b) of FIG. 1 shows an underexposed image. By comparing Figure 1 (a) of FIG. 1 and Figure 1 (b) of FIG. 1, it can be very intuitively felt that Figure 1 the normally exposed image shown in (a) of FIG. 1 has an overall normal brightness, Figure 1 the underexposed image shown in (b) of FIG. 1, especially the light background 1, is in a dark state.

[0131] Overexposure refers to excessive exposure of an image, resulting in an overall bright image and loss of details in the highlight part. In the case of overexposure, the bright part of the image is too bright, showing a dead white, and the level and texture cannot be distinguished, while the dark part is relatively normal, but the overall brightness and contrast of the image are unbalanced.

[0132] Exemplarily, Figure 2 A comparison diagram of a normally exposed image and an overexposed image is given.

[0133] Figure 2 A comparison diagram of a normally exposed image and an overexposed image of a bracelet placed on a black table (dark background 2) is given. Among them, Figure 2 (a) of FIG. 2 shows a normally exposed image, Figure 2 (b) of FIG. 2 shows an overexposed image. By comparing Figure 2 (a) of FIG. 2 and Figure 2 (b) of FIG. 2, it can be very intuitively felt that Figure 2the overall normal brightness in the normal exposure image shown in (a) of FIG. 1, Figure 2 the over-bright state of the dark background 2 in the overexposure image shown in (b) of FIG. 1.

[0134] In an image processing scenario or some shooting scenario, the exposure state of an image needs to be quantified (i.e., evaluated) to determine whether the image has an underexposure or overexposure problem.

[0135] A common image exposure quantification method is a subjective evaluation method. The subjective evaluation method is to use a person as an observer to make a subjective judgment on the exposure state of an image, thereby obtaining a quantification result of the image. The subjective evaluation method can directly reflect the visual perception of the image to a person and is the most intuitive exposure quantification method. However, there can be a large difference between different observers in the subjective evaluation method, and the subjective evaluation method has strong subjectivity and subjective ambiguity. Moreover, the entire process of the subjective evaluation method relies on manual work, and the evaluation process is time-consuming and labor-intensive.

[0136] Exemplarily, Figure 3 A schematic diagram of suitable exposure images selected by different observers in the subjective evaluation method is given. Figure 3 In the diagram, Figure 3 Image 1 shown in (a) of FIG. 1 is a suitable exposure image (same as a normal exposure image) selected by observer A; Figure 3 Image 2 shown in (b) of FIG. 1 is a suitable exposure image (same as a normal exposure image) selected by observer B; Figure 3 Image 3 shown in (c) of FIG. 1 is a suitable exposure image (same as a normal exposure image) selected by observer C. It can be felt from the diagram that Figure 3 In the diagram, Figure 3 (a) of FIG. 1, Figure 3 (b) of FIG. 1, and Figure 3 (c) of FIG. 1 are images in different exposure states. That is, for the images of the same shooting content, there is a large difference in the evaluation of “normal exposure” by observer A, observer B, and observer C, the exposure evaluation has subjectivity, and the final exposure evaluation of the images is not accurate.

[0137] For exposure quantification of an image, the average of exposure quantification values of multiple observers can be calculated to obtain a final exposure quantification result of the image.

[0138] Exemplarily, the calculation expression of an image exposure quantification result MOS is as follows:

[0139] Expression (1).

[0140] wherein, is an exposure quantification value of an i th observer, and i has a value of 1, 2, … k; for the i-th observer.

[0141] In addition to the subjective evaluation method described above, a common image exposure quantification method is an objective evaluation method. The objective evaluation method is divided into two types of methods, namely, a reference method and a non-reference method. The reference method refers to the need for a reference image, that is, a reference image is needed in the process of image exposure quantification to compare the image exposure state, so as to obtain the exposure quantification result of the image to be evaluated. Although this method can objectively obtain the exposure quantification result of the image to be evaluated, it is troublesome to find a reference image in the process of image exposure quantification. The non-reference method refers to an objective exposure quantification without a reference image. For example, the non-reference image exposure quantification method includes exposure prediction based on some preset statistical indicators, so as to obtain the exposure quantification result of the image to be evaluated. The statistical indicators can include the 18-degree gray business convention indicators. However, this method may have exposure prediction errors for some special scenes, and thus needs to be repeatedly corrected in the exposure quantification process, resulting in inaccurate and time-consuming exposure quantification results. For example, the special scenes include scenes containing large-area light content such as snow scenes, or scenes containing large-area dark content. For these special scenes, the non-reference image exposure quantification method may have exposure prediction errors, and thus needs secondary exposure prediction correction, which has certain limitations.

[0142] For example, Figure 4 A before-and-after comparison chart of the corrected exposure of a snow scene is given. Figure 4 As shown in (a) of FIG. 1, the image obtained by the non-reference image exposure quantification method is considered to be a properly exposed image. Obviously, for the image with a large-area snow scene, the image is actually too dark. Figure 4 As shown in (a) of FIG. 1, the image is actually too dark. It is assumed that Figure 4 As shown in (a) of FIG. 1, the image 1 has an exposure compensation of 0 and an aperture value of F18 and a shutter speed of 1 / 250 s. The exposure parameters of the shooting scene are corrected, the exposure compensation is increased from 0 to 4 / 3, the aperture value is kept unchanged at F18, and the shutter speed is reduced to 1 / 100 s, to obtain the image 2 after exposure compensation. Figure 4 As shown in (b) of FIG. 1, the image 2. Comparing Figure 4 (a) of FIG. 1 and Figure 4 (b) of FIG. 1, it can be intuitively felt that the image 1 before exposure correction shown in (a) of FIG. 1 is too dark. Figure 4 The snow scene in the image 1 shown in (a) of FIG. 1 is too dark, and the snow scene in the image 2 after exposure correction shown in (b) of FIG. 1 is properly bright. Figure 4 The snow scene in the image 1 shown in (a) of FIG. 1 is too dark, and the snow scene in the image 2 after exposure correction shown in (b) of FIG. 1 is properly bright.

[0143] In addition, the image exposure quantification method of the non-parametric class also includes other methods. For example, the exposure of the image is quantified through a quality evaluation model.

[0144] On the one hand, the quality evaluation model usually needs to be manually annotated, involves manual annotation, and has various annotation quantification methods and quantification values, so it is difficult to obtain a unified measurement standard.

[0145] The following Table 1 gives a comparison of the image subjective evaluation scale.

[0146] Table 1

[0147]

[0148] On the other hand, many quality evaluation models are not designed for image exposure quantification. In the quality evaluation process based on the image quality assessment (IQA) dataset, the IQA dataset usually involves different degradation types to simulate various quality problems that images may encounter in the real world. For example, common degradation types include blur, noise, compression distortion, and contrast distortion. If the IQA dataset is used for image exposure quantification, the quality evaluation model and the IQA dataset need to be fine-tuned or migrated, which makes the image exposure quantification process complex.

[0149] Figure 5 A common quality evaluation model training approach is given.

[0150] The quality evaluation dataset (IQA dataset) can include multiple sample images. The IQA dataset is input into an initial quality evaluation model as a training dataset. The initial quality evaluation model performs quality evaluation based on one or more quantification strategies to obtain an initial predicted quality quantification value corresponding to the training dataset. A Pearson linear correlation coefficient (PLCC) is calculated based on the initial predicted quality quantification value of the input dataset. The PLCC is used to measure the linear relationship between the initial predicted quality quantification value and the subjective quality quantification value. The quality evaluation model is fine-tuned based on the PLCC and a fine-tuning dataset to obtain a fine-tuned quality evaluation model. The fine-tuning dataset includes multiple images in different exposure states. The fine-tuned quality evaluation model can be used for exposure quantification. Thus, the training dataset is input into the fine-tuned quality evaluation model for inference to obtain a final exposure quantification result.

[0151] The training process of the whole quality evaluation model needs to obtain images in different exposure states to fine-tune the quality evaluation model, and the obtaining of the images in different exposure states needs a certain time and procedure, the whole process is complicated and makes the training process of the quality evaluation model inefficient.

[0152] The embodiment of the present application provides an image exposure quantification method, which obtains a semantic similarity value between an input image and a preset prompt word through a trained image-text comparison model, wherein the preset prompt word includes a field describing different exposure states. The exposure state label of the input image can be determined based on the semantic similarity value. The image-text comparison model can learn the semantic relationship between the preset prompt word (text) and the image, thereby obtaining the semantic similarity value of the image and the preset prompt word. The exposure state label of the image is obtained based on the semantic similarity value, which is objective and accurate. The subjective ambiguity of manual evaluation is avoided, and the whole process does not need to find reference benchmark images or develop complex quantification strategies, which reduces the computing power while ensuring the accuracy of exposure quantification.

[0153] Further, the training process of the trained image-text comparison model used in the image exposure quantification method provided by the embodiment of the present application can be performed by sample images and training prompt words, so that the image-text comparison model has more accurate understanding of the exposure of the image, and the similarity between the output image and the training prompt word is more accurate. Compared with the conventional image exposure quantification method based on the quality evaluation model, the embodiment of the present application only needs to adjust the model parameters (such as model weights) and the training prompt words in the training process of the image-text comparison model, without developing complex quantification strategies, and without fine-tuning the model based on a large number of images in different exposure states. The model training is simpler and more efficient while ensuring the accuracy of the output of the image-text comparison model.

[0154] The application scenario of the image exposure quantification method provided by the embodiment of the present application can be a pre-processing scenario of an image. For example, in an electronic device with a camera, the electronic device displays a preview image when starting the camera. The electronic device can perform exposure quantification based on the preview image, and determine whether exposure compensation is needed for the preview image based on the exposure state label of the preview image. In the case where the shooting scene does not change, the corresponding exposure compensation processing can also be performed on the shooting image when a shooting operation is received.

[0155] The application scenario of the image exposure quantification method provided by the embodiment of the present application can also be a post-processing scenario of an image. For example, in some electronic devices with image processing capability, the electronic device can perform exposure quantification on a batch of to-be-processed images to obtain the exposure state label of the to-be-processed images. The embodiment does not limit the specific application scenario of the image exposure quantification method.

[0156] The image exposure quantification method provided by the embodiments of the present application can be applied to an electronic device with image processing capability. Optionally, the electronic device can have a shooting capability. Exemplarily, the electronic device can be a portable computer (such as a mobile phone), a tablet computer, a notebook computer, a personal computer (PC), a wearable electronic device (such as a smart watch), an augmented reality (AR) \ virtual reality (VR) device, a vehicle-mounted computer, a server, and the like. The specific form of the electronic device is not specially limited in the following embodiments.

[0157] Figure 6 A structural schematic diagram of the electronic device 100 is shown.

[0158] The electronic device 100 can include a processor 110, an external memory interface 120, an internal memory 121, a universal serial bus (USB) interface 130, a charging management module 140, a power management module 141, a battery 142, an antenna 1, an antenna 2, a mobile communication module 150, a wireless communication module 160, a sensor module 180, a camera 193, and a display screen 194, and the like.

[0159] It can be understood that the structure shown in the embodiments of the present application does not constitute a specific limitation on the electronic device 100. In other embodiments of the present application, the electronic device 100 can include more or fewer components than shown, or combine certain components, or split certain components, or different arrangement of components. The components shown can be implemented in hardware, software, or a combination of software and hardware.

[0160] The processor 110 can include one or more processing units, for example: the processor 110 can include an application processor (AP), a modem processor, a graphics processing unit (GPU), an image signal processor (ISP), a controller, a memory, a video codec, a digital signal processor (DSP), a baseband processor, and / or a neural-network processing unit (NPU), and the like. Different processing units can be independent devices, or can be integrated in one or more processors.

[0161] The controller can be the nerve center and command center of the electronic device 100. The controller can generate operation control signals according to instruction operation codes and timing signals, and complete the control of instruction fetching and instruction execution.

[0162] The processor 110 can also be provided with a memory for storing instructions and data. In some embodiments, the memory in the processor 110 is a cache memory. The memory can store instructions or data that have just been used or are recycled by the processor 110. If the processor 110 needs to use the instructions or data again, it can be directly called from the memory. This avoids repeated access and reduces the waiting time of the processor 110, thereby improving the efficiency of the system.

[0163] In this embodiment, the processor 110 can serve as the execution subject of the image exposure quantization method, and be configured to perform exposure quantization on an input image to obtain an exposure quantization result of the input image.

[0164] In some embodiments, the processor 110 can include one or more interfaces. The interfaces can include an inter-integrated circuit (I2C) interface, an inter-integrated circuit sound (I2S) interface, a pulse code modulation (PCM) interface, a universal asynchronous receiver / transmitter (UART) interface, a mobile industry processor interface (MIPI), a general-purpose input / output (GPIO) interface, a subscriber identity module (SIM) interface, and / or a universal serial bus (USB) interface, etc.

[0165] It can be understood that the interface connection relationship between the modules shown in the embodiments of the present application is only illustrative and does not constitute a limitation on the structure of the electronic device 100. In other embodiments of the present application, the electronic device 100 can also use different interface connection methods or combinations of multiple interface connection methods in the above embodiments.

[0166] The charging management module 140 is configured to receive charging input from a charger. The charger can be a wireless charger or a wired charger. In some embodiments with wired charging, the charging management module 140 can receive charging input from a wired charger through the USB interface 130. In some embodiments with wireless charging, the charging management module 140 can receive wireless charging input through a wireless charging coil of the electronic device 100. The charging management module 140 can supply power to the electronic device while charging the battery 142.

[0167] The power management module 141 is configured to connect the battery 142, the charging management module 140, and the processor 110. The power management module 141 receives input from the battery 142 and / or the charging management module 140 to supply power to the processor 110, the internal memory 121, the external memory, the display screen 194, the camera 193, and the wireless communication module 160. The power management module 141 can also be configured to monitor parameters such as battery capacity, battery cycle count, battery health status (leakage, impedance), and the like. In some other embodiments, the power management module 141 can also be disposed in the processor 110. In some other embodiments, the power management module 141 and the charging management module 140 can also be disposed in the same device.

[0168] The wireless communication function of the electronic device 100 can be implemented through the antenna 1, the antenna 2, the mobile communication module 150, the wireless communication module 160, the modem processor, and the baseband processor.

[0169] The electronic device 100 can implement the display function through the GPU, the display screen 194, and the application processor. The GPU is a microprocessor for image processing, which is connected to the display screen 194 and the application processor. The GPU is configured to perform mathematical and geometric calculations for graphics rendering. The processor 110 can include one or more GPUs, which execute program instructions to generate or change display information.

[0170] The display screen 194 is configured to display images, videos, and the like. The display screen 194 includes a display panel. The display panel can be a liquid crystal display (LCD), an organic light-emitting diode (OLED), an active-matrix organic light-emitting diode (AMOLED), a flex light-emitting diode (FLED), a Miniled, a MicroLed, a Micro-oLed, a quantum dot light emitting diode (QLED), or the like. In some embodiments, the electronic device 100 can include one or N display screens 194, where N is a positive integer greater than 1.

[0171] In some embodiments, the electronic device can display, through the display screen 194, a preview image after exposure compensation based on the exposure state label of the preview image, or based on a captured image after exposure compensation, and the like.

[0172] The electronic device 100 can implement the photographing function through the ISP, the camera 193, the video codec, the GPU, the display screen 194, and the application processor, and the like.

[0173] The ISP is configured to process data fed back by the camera 193. For example, when taking a photo, the shutter is opened, light is transmitted to the camera photosensitive element through the lens, the light signal is converted into an electrical signal, and the camera photosensitive element transmits the electrical signal to the ISP for processing to convert it into an image visible to the naked eye. The ISP can also optimize the algorithm for noise, brightness, and skin color of the image. The ISP can also optimize the exposure, color temperature, and other parameters of the shooting scene. In some embodiments, the ISP can be disposed in the camera 193.

[0174] The camera 193 is configured to capture still images or videos. An object projects an optical image through a lens to a photosensitive element. The photosensitive element can be a charge coupled device (CCD) or a complementary metal-oxide-semiconductor (CMOS) phototransistor. The photosensitive element converts the optical signal into an electrical signal, which is then transmitted to an ISP to be converted into a digital image signal. The ISP outputs the digital image signal to a DSP for processing. The DSP converts the digital image signal into an image signal in a standard format, such as RGB, YUV, or the like. In some embodiments, the electronic device 100 can include one or N cameras 193, where N is a positive integer greater than 1.

[0175] After the exposure quantization, when the electronic device captures a preview image or a shot image, the electronic device can perform exposure compensation on the preview image or the shot image based on the exposure state label, such as reducing the opening speed, or increasing the exposure compensation parameter, and the like.

[0176] The digital signal processor is configured to process digital signals, in addition to processing digital image signals, it can also process other digital signals. For example, when the electronic device 100 selects a frequency point, the digital signal processor is configured to perform Fourier transform on the frequency point energy, and the like.

[0177] The external memory interface 120 can be configured to connect an external memory card, such as a Micro SD card, to expand the storage capacity of the electronic device 100. The external memory card communicates with the processor 110 through the external memory interface 120 to implement a data storage function. For example, files such as music and videos are stored in the external memory card.

[0178] The internal memory 121 can be configured to store computer executable program codes including instructions. The processor 110 executes various functional applications and data processing of the electronic device 100 by running the instructions stored in the internal memory 121. The internal memory 121 can include a program storage area and a data storage area. The program storage area can store an operating system, at least one application program required by a function (such as a sound playing function, an image playing function, and the like), and the like. The data storage area can store data created during use of the electronic device 100 (such as audio data, a phonebook, and the like), and the like. In addition, the internal memory 121 can include a high-speed random access memory, and can also include a non-volatile memory, such as at least one magnetic disk storage device, a flash memory device, a universal flash storage (UFS), and the like.

[0179] Embodiments of the present application do not limit the specific structure of the electronic device.

[0180] For the method of quantifying image exposure, in addition to the common image processing related method, the image exposure can also be quantified by the semantic understanding of the image and the text, through the semantic understanding and semantic matching of the image exposure and the exposure prompt word. In the embodiment of the present application, the characteristics of semantic features can be learned by using a large language model (LLM), and the image feature learning ability can be further expanded on the basis of the text feature learning ability provided by the LLM. The image exposure state based on the exposure state prompt word is understood by using the image-text feature learning ability of the image-text comparison model, and the image exposure quantification result is obtained.

[0181] Exemplarily, Figure 7 A set of images in different exposure states are given. Among them, Figure 7 Including 32 images of the same shooting content in different exposure states, it can be seen intuitively that the 32 images are in the order of image 1 to image 32, and the exposure presents an increasing state. Based on different exposure state definition prompt words, for example, the prompt words can include underexposed, well exposed and overexposed. The prompt words and sample images are input into the image-text comparison model to calculate the similarity value between the image and the text.

[0182] Figure 8 A broken line diagram of the similarity value between different prompt words and images is given. Among them, the horizontal axis of the broken line diagram is the image number, from 1 to 32 including 32 images. The vertical axis of the broken line diagram is the similarity value.

[0183] From Figure 8 It can be seen that as the image number increases, the similarity between the image and the prompt word "overexposed" gradually increases. The similarity between the image and the prompt word "underexposed" gradually decreases. By observing Figure 7 It can be seen that as the image number increases, the image exposure increases, and the image gradually changes from underexposure to overexposure. That is to say, the image-text comparison model can learn that as the image number increases, the image exposure also increases.

[0184] Based on the image-text comparison model, the information related to the image exposure state can be captured, and the semantic features learned by the image-text comparison model can be used to infer the quantitative results related to the image quality from the image-text comparison results, including the exposure quantification result. Based on this, the image-text comparison model can be trained to obtain the similarity value between the image and the text describing the exposure state, so as to obtain the exposure state label of the image based on the similarity value.

[0185] In one embodiment, Figure 9 A diagram of the training path for an image-text comparison model is given, which includes:

[0186] S101: The electronic device obtains a sample data set.

[0187] The sample data set may include multiple sample image groups. Each sample image group includes at least two images with different exposure states and the same shooting content. Figure 9 As shown, the image numbered 1 and the image numbered 2 are a sample image group, the image numbered 3 and the image numbered 4 are a sample image group, the image numbered 5 and the image numbered 6 are a sample image group, and the image numbered 7 and the image numbered 8 are a sample image group.

[0188] In some embodiments, to ensure more accurate model training, each sample image group may include at least three images with the same content but different exposure states. That is, at least one image is required for each of the following exposure states: underexposure, normal exposure, and overexposure.

[0189] Each sample image in the sample data set may be an image generated by image processing, or may be an image captured and collected under different exposures, which is not limited in this embodiment.

[0190] S102: The electronic device obtains a first training prompt word.

[0191] The first training prompt word can be text used to describe the exposure state of an image. For example, the first training prompt word can be a word representing underexposure, normal exposure, or overexposure. For example, the first training prompt word can include text containing the keywords underexposed, normal exposure, and overexposed. The first training prompt word corresponds to the prompt word of the sample dataset in this embodiment.

[0192] For example, the prompt words may be "this photo is underexposed", "this foreground of photo is underexposed", "this subject of photo is underexposed", "this photo is wellexposed", "this photo is overexposed", etc.

[0193] The electronic device can define the prompt word based on an actual scene, and the description range of all prompt words needs to at least include descriptions of the three states of underexposure, normal exposure, and overexposure.

[0194] In some other possible embodiments, the characterization of the exposure state of the image can be understood as the brightness of the image to some extent. The prompt word can also be other text characterizing the exposure state. For example, the prompt word can be "this image is too bright", "this image is too dark", "the brightness of this image is appropriate", and the like.

[0195] The specific content and quantity of the prompt word are not limited in the embodiments of the present application.

[0196] In S103, the electronic device inputs the sample data set and the first training prompt word into the image-text contrast model for inference to obtain a similarity value of the sample image and the first training prompt word.

[0197] The image-text contrast model can be a contrastive language-image pre-training (CLIP) model. The CLIP model can learn the semantic relationship between images and texts through natural language supervision. That is, the CLIP model can calculate the similarity value of each sample image and the first training prompt word, and the similarity value can reflect the exposure state of the sample image.

[0198] In the embodiment, the electronic device inputs the sample data set and the first training prompt word into the CLIP model to obtain the similarity value of the sample image in the sample data set and the first training prompt word. In the case where a plurality of first training prompt words are included, the CLIP model outputs the similarity value of the sample image and each first training prompt word.

[0199] In some possible embodiments, the CLIP model can output the cosine similarity of the sample image and the preset prompt word. For example, the first training prompt word includes underexposed, well exposed, and overexposed. Then the cosine similarity of the sample image and the preset prompt word can be represented as:

[0200] Formula (2).

[0201] Wherein, may represent the cosine similarity (i.e., the similarity value) of the sample image and the first training prompt word 1 (e.g., underexposed), ​The cosine similarity (i.e., a similarity value) of the sample image to the first training hint 2 (e.g., wellexposed) can be represented as: The cosine similarity (i.e., a similarity value) of the sample image to the first training hint 3 (e.g., overexposed) can be represented as:

[0202] The cosine similarity can be a specific numerical value. For example, the similarity value of the sample image 1 to the first training hint 1 (underexposed) is 80%, the similarity value of the sample image 1 to the first training hint 2 (wellexposed) is 10%, and the similarity value of the sample image 1 to the first training hint 3 (overexposed) is 8%. In this way, the similarity value of the sample image corresponding to the first training hint is obtained.

[0203] In S104, the electronic device determines the training label of the sample image based on the similarity value of the sample image.

[0204] In some embodiments, the electronic device can take the maximum cosine similarity as the exposure quantization value, and determine the training label of the sample image based on the first training hint corresponding to the maximum cosine similarity. The training label refers to the exposure quantization label of the sample image output by the image-text contrast model.

[0205] In a feasible manner, the maximum cosine similarity is taken as the exposure quantization value which can be represented as:

[0206] Equation (3).

[0207] For example, in the above S103, the maximum cosine similarity is 80%, and the cosine similarity of the first training hint 1 is taken as the exposure quantization value of the sample image.

[0208] In a feasible manner, the training label of the sample image can be determined based on the first training hint corresponding to the maximum cosine similarity, which can be represented as:

[0209] Equation (4).

[0210] For example, in the above S103, the maximum cosine similarity is 80%, and the training label (i.e., index) of the sample image is determined based on the first training hint 1 (underexposed) as “underexposed”.

[0211] S105, the electronic device calculates a loss based on the training label of the sample image and the reference label corresponding to the sample image, iteratively trains the image-text contrast model based on the loss value, until a training condition is met, and obtains the trained image-text contrast model.

[0212] The reference label refers to a label that can represent the actual exposure state of the sample image. The reference label can be a manually annotated label or a label determined based on the exposure parameter by the electronic device.

[0213] In this embodiment, the electronic device can calculate the loss (such as a correlation loss) between the training label of the sample image and the reference label corresponding to the sample image, iteratively train the image-text contrast model based on the loss, adjust the model parameters and training prompt words, until the training condition is met, thereby obtaining the trained image-text contrast model. The training condition can be that the loss gradually approaches 0, or the number of iterations meets a number threshold. Model training can make the image-text contrast model more accurate in understanding the exposure of the image, and make the similarity between the output sample image and the training prompt word more accurate, so that the exposure state label of the image obtained based on the trained image-text contrast model is more accurate.

[0214] The image exposure quantification method provided in this embodiment does not need to develop a complex quantification strategy, and does not need to fine-tune the model based on the reference image, but the trained image-text contrast model can still guarantee the accuracy of the understanding of the exposure state of the image. That is, the image exposure quantification method provided in this embodiment ensures the accuracy of the output result of the image-text contrast model, and only needs to adjust the model parameters (such as model weights) and the training prompt word in the model training process, without fine-tuning the entire model. The model training is simpler and more efficient.

[0215] Correspondingly, the process of quantifying the image exposure of the input image based on the trained image-text contrast model can refer to Figure 10 A flowchart of an image exposure quantification method is given, which includes:

[0216] S201, the electronic device obtains an input image.

[0217] The input image can be a to-be-quantified image in different scenes. For example, in a shooting scene, the input image can be a preview image or a shooting image. In an image processing scene, the input image can be an image that needs to be quantified.

[0218] S202, the electronic device inputs the input image into the image-text contrast model, and obtains a similarity value corresponding to a preset prompt word of the input image.

[0219] The electronic device stores a preset prompt word. The preset prompt word can be the same as the first training prompt word. The input image is input into the image-text comparison model. The image-text comparison model can obtain the preset prompt word, calculate the similarity between the input image and the preset prompt word, and obtain a plurality of similarity values of the input image corresponding to the preset prompt word. The specific manner of calculating the similarity value can refer to the embodiment provided in S103 in the model training embodiment, and details are not repeated in this embodiment.

[0220] In S203, the electronic device determines the exposure state label corresponding to the input image based on the similarity value of the input image corresponding to the preset prompt word.

[0221] The specific manner of determining the exposure state label corresponding to the input image can refer to the embodiment provided in S104 in the model training embodiment, and details are not repeated in this embodiment.

[0222] In this embodiment, the trained image-text comparison model can realize semantic understanding of the relevance of the image and the preset prompt word. The method of quantifying the image exposure based on the image-text comparison model avoids the problem of inaccurate quantification and non-uniform standard caused by subjective ambiguity in manual quantitative evaluation, and improves the accuracy of image exposure quantification.

[0223] In some embodiments, the image that needs to be quantified for exposure may be an image with a large area of light color (such as white) content or an image with a large area of dark color (such as black) content as shown in Figure 4 For these special scene images, some image preprocessing can be performed on the input image / sample image before inputting it into the image-text comparison model for similarity value calculation to obtain a more accurate exposure state label of the image.

[0224] In some feasible embodiments, in order to avoid the interference of a large area of light color (such as white) or dark color (such as black) background on image exposure quantification, subject detection can be performed on the subject of the image to obtain a local image corresponding to the subject. That is, the image preprocessing in this embodiment can be subject detection on the subject in the global image to obtain the subject image.

[0225] Therefore, the electronic device can input the original image (i.e., the global image) and the prompt word into the image-text comparison model to obtain the similarity value of the global image and the prompt word. The subject image (i.e., the local image) and the prompt word are input into the image-text comparison model to obtain the similarity value of the subject image and the prompt word. The comprehensive similarity value of the image is obtained by combining the similarity value of the global image and the prompt word and the similarity value of the subject image and the prompt word, and the exposure state label of the image is determined based on the comprehensive similarity value.

[0226] In some embodiments, for example, Figure 11A training path diagram of another image-text contrast model is given, which includes:

[0227] In S301, the electronic device acquires a sample data set.

[0228] The sample data set can include a plurality of sample image groups. Each sample image group includes at least two images with different exposure states and the same shooting content. The sample images in the sample data set can also be referred to as global images.

[0229] The specific manner of acquiring the sample data set can refer to the embodiments provided in S101, and the embodiments will not be described herein.

[0230] In S302, the electronic device acquires a first training prompt.

[0231] The first training prompt can be a text used to describe the exposure state of the global image. The specific manner of acquiring the first training prompt can refer to the embodiments provided in S101, and the embodiments will not be described herein.

[0232] In this embodiment, exemplary first training prompts can be "this photo is underexposed", "this foreground of photo is underexposed", "this subject of photo is underexposed", "this photo is well exposed", "this photo is overexposed", and the like.

[0233] In S303, the electronic device acquires a subject image corresponding to the sample image.

[0234] In this embodiment, the electronic device can perform saliency model inference on the sample images in the sample data set to acquire the subject images corresponding to the sample images. The saliency model inference refers to analyzing and explaining important features or regions (such as the subject region) in the input data (such as the sample image) through the saliency model. The saliency model inference in this embodiment can refer to the saliency model inference in the conventional manner, which will not be described herein.

[0235] In this embodiment, the subject image can also be referred to as a local image relative to the entire sample image.

[0236] For example, Figure 11As shown, after significant model inference, the No. 1 sample image corresponding to the No. 1 subject image, the No. 2 sample image corresponding to the No. 2 subject image, the No. 3 sample image corresponding to the No. 3 subject image, the No. 4 sample image corresponding to the No. 4 subject image, the No. 5 sample image corresponding to the No. 5 subject image, the No. 6 sample image corresponding to the No. 6 subject image, the No. 7 sample image corresponding to the No. 7 subject image, and the No. 8 sample image corresponding to the No. 8 subject image are obtained.

[0237] In S304, the electronic device obtains a second training prompt.

[0238] The second training prompt can be text used to describe the exposure state of a local image. The specific manner of obtaining the second training prompt can refer to the embodiments provided in S101 above, and details are not described herein.

[0239] In this embodiment, the second training prompt can be the same as or different from the first training prompt. The same includes the same semantics or the same description.

[0240] In this embodiment, the second prompt can be, for example, "this photo is underexposed", "this subject of photo is underexposed", "this photo is well exposed", "this photo is overexposed", and the like.

[0241] In S305, the electronic device inputs the sample data set and the first training prompt into the image-text comparison model for inference to obtain a first similarity value of the sample image and the first training prompt.

[0242] In this embodiment, the electronic device inputs the sample data set and the first training prompt into the CLIP model to obtain the similarity value of the sample image in the sample data set and the first training prompt. In the case of containing multiple first training prompts, the CLIP model outputs the similarity value of the sample image and each first training prompt.

[0243] The specific method of calculating the first similarity value can refer to the embodiments provided in S103 above. For example, in this embodiment, the obtained first similarity value is .

[0244] S306, the electronic device inputs the subject image and the second training prompt word into the text-image comparison model for inference to obtain a second similarity value of the subject image and the second training prompt word.

[0245] In this embodiment, the electronic device inputs the subject image and the second training prompt word into the CLIP model to obtain a similarity value of the subject image and the second training prompt word. In the case of containing multiple second training prompt words, the CLIP model outputs a similarity value of the sample image and each second training prompt word.

[0246] The specific method of calculating the second similarity value can refer to the embodiments provided in S103 above. For example, in this embodiment, the obtained second similarity value is .

[0247] S307, the electronic device obtains a third similarity value corresponding to the sample image according to the first similarity value and the second similarity value.

[0248] In this embodiment, the electronic device can define a first weight corresponding to the first similarity value and a second weight corresponding to the second similarity value.

[0249] The electronic device performs weighted calculation based on the first similarity value, the first weight, the second similarity value, and the second weight to obtain the third similarity value of the sample image.

[0250] In this embodiment, the first similarity used for weighted calculation refers to the maximum value in the first similarity, and the second similarity used for weighted calculation refers to the maximum value in the second similarity.

[0251] The third similarity value y can be obtained in the following manner:

[0252] Equation (5).

[0253] wherein, is the first weight, is the maximum first similarity value corresponding to the sample image; is the second weight, is the maximum second similarity value corresponding to the local image.

[0254] S308, the electronic device determines the training label of the sample image based on the third similarity value of the sample image.

[0255] Generally, the similarity of the sample image to different first training prompt words has the same trend as the similarity of the local image to different second training prompt words. That is, the target first training prompt word corresponding to the maximum first similarity value of the sample image and the target second training prompt word corresponding to the maximum second similarity value of the local image have consistency in representing the exposure state.

[0256] Based on this, the third similarity value after the weighting calculation can represent the corrected similarity value of the sample image to the target first training prompt word and the target second training prompt word.

[0257] Based on the third similarity value, the corresponding target first training prompt word and the target second training prompt word, the training label of the sample image can be determined.

[0258] S309, the electronic device calculates the loss based on the training label of the sample image and the reference label corresponding to the sample image, iteratively trains the image-text contrast model based on the loss value, until the training condition is met, and obtains the trained image-text contrast model.

[0259] In this embodiment, the training process of the image-text contrast model can refer to the embodiments provided in S105 described above, and this embodiment will not be repeated here.

[0260] The image exposure quantification method provided in this embodiment does not need to develop a complex quantification strategy, and does not need to fine-tune the model based on the reference image, but the trained image-text contrast model can still guarantee the accuracy of the understanding of the exposure state of the image. That is, the image exposure quantification method provided in this embodiment only needs to adjust the model parameters (such as model weights) and training prompt words, and does not need to obtain image data of different exposure states for model fine-tuning. While ensuring the accuracy of the output results of the image-text contrast model, the model training process is simpler and the model training efficiency is higher. Moreover, in this embodiment, the correlation learning of the local image and the training prompt word is introduced, so that the image-text contrast model can obtain more accurate exposure quantification results for the entire image by exposure quantification of the main body for some images containing special scenes, such as images with large-area light background or large-area dark background, so that the image exposure quantification is more accurate.

[0261] Correspondingly, the process of image exposure quantification of the input image based on the trained image-text contrast model can refer to the flowchart of another image exposure quantification method provided in Figure 12 The flowchart of another image exposure quantification method provided in

[0262] S401, the electronic device obtains an input image.

[0263] The method provided in the foregoing embodiment S201 can be referred to specifically, and no redundant description is given in this embodiment.

[0264] S402, the electronic device performs significant model inference on the input image to obtain a subject image corresponding to the input image.

[0265] The method provided in the foregoing embodiment S303 can be referred to specifically, and no redundant description is given in this embodiment.

[0266] S403, the electronic device inputs the input image into a graph-text comparison model to obtain a first similarity value corresponding to a preset prompt word of the input image.

[0267] The preset prompt word is stored in the electronic device. The preset prompt word can include a first training prompt word and a second training prompt word. When the input image is input into the graph-text comparison model, the graph-text comparison model can obtain the preset prompt word, calculate the similarity between the input image and the preset prompt word, and thus obtain multiple first similarity values corresponding to the preset prompt word of the input image. The specific manner of calculating the first similarity value can refer to the embodiment provided in S305 in the model training embodiment, and no redundant description is given in this embodiment.

[0268] S404, the electronic device inputs the subject image into the graph-text comparison model to obtain a second similarity value corresponding to the preset prompt word of the subject image.

[0269] When the input image is input into the graph-text comparison model, the graph-text comparison model can obtain the preset prompt word, calculate the similarity between the input image and the preset prompt word, and thus obtain multiple second similarity values corresponding to the preset prompt word of the input image. The specific manner of calculating the second similarity value can refer to the embodiment provided in S306 in the model training embodiment, and no redundant description is given in this embodiment.

[0270] S405, the electronic device calculates a third similarity value corresponding to the input image based on the first similarity value and the second similarity value.

[0271] The specific manner of calculating the third similarity value can refer to the embodiment provided in S307 in the model training embodiment, and no redundant description is given in this embodiment.

[0272] S406, the electronic device determines an exposure state label corresponding to the input image based on the third similarity value.

[0273] The specific manner of determining the exposure state label corresponding to the input image based on the third similarity value can refer to the embodiment provided in S308 in the model training embodiment, and no redundant description is given in this embodiment.

[0274] In this embodiment, the trained image-text comparison model can realize semantic understanding of the relevance of the global image of the input image to the preset prompt word, and semantic understanding of the relevance of the subject image corresponding to the input image to the preset prompt word. For an input image containing some special scenes, performing semantic understanding of the relevance of the subject image to the preset prompt word can better determine the actual exposure state of the input image, so that the exposure quantization label of the final obtained input image is more accurate.

[0275] The above Figure 10 And Figure 12 The image exposure quantization method applied by the model given above can be understood as a scene in which one image corresponds to multiple prompt words, and the exposure state of the one image is determined by the similarity between the one image and the multiple prompt words.

[0276] In some other feasible embodiments, the trained image-text comparison model can provide semantic understanding capability between images and preset prompt words. Based on this capability, target images that meet a target prompt word can also be determined from a group of images. That is, the image exposure quantization method provided by the embodiments of the present application can also be used to query target images that match the target prompt word. For example, a group of images includes multiple images with different exposure states, and from the multiple images with different exposure states, a target image with "appropriate exposure" is queried.

[0277] In some embodiments, Figure 13 Another flowchart of an image exposure quantization method is given, which includes:

[0278] S501, the electronic device obtains multiple input images.

[0279] Among them, the multiple input images include images with different exposure states, such as at least two of underexposed images, normally exposed images, and overexposed images. For example, the multiple input images can include images numbered 1-5 with different exposure states as shown in Figure 13

[0280] S502, the electronic device performs significant model reasoning on each input image to obtain a subject image corresponding to each input image.

[0281] For details, reference can be made to the method provided in the above embodiment S303, and the present embodiment will not be repeated.

[0282] In this embodiment, subject images numbered 1'-5' corresponding to images numbered 1-5 can be obtained.

[0283] S503, the electronic device inputs each input image into the image-text comparison model to obtain a fourth similarity value corresponding to each input image to a target prompt word.

[0284] ​The target prompt word can include a prompt word of "normal exposure".

[0285] The embodiment aims to determine a target image meeting the prompt word from multiple input images based on a similarity value between the input images and the prompt word.

[0286] Here, each input image is input into the image-text contrast model, and the image-text contrast model can calculate a similarity between each input image and the target prompt word based on the target prompt word, thereby obtaining a fourth similarity value corresponding to the target prompt word for each input image.

[0287] S504, the electronic device inputs each subject image into the image-text contrast model to obtain a fifth similarity value corresponding to the target prompt word for each subject image.

[0288] Here, each subject image is input into the image-text contrast model, and the image-text contrast model can calculate a similarity between each subject image and the target prompt word based on the target prompt word, thereby obtaining a fourth similarity value corresponding to the target prompt word for each subject image.

[0289] S505, the electronic device calculates a sixth similarity value corresponding to the input image based on the fourth similarity value of the input image and the fifth similarity value of the subject image corresponding to the input image.

[0290] Here, the electronic device can perform weighted calculation on the fourth similarity value and the fifth similarity value to obtain the sixth similarity value. Alternatively, the electronic device can take an average of the fourth similarity value and the fifth similarity value as the sixth similarity value. Alternatively, the electronic device can take a larger similarity value between the fourth similarity value and the fifth similarity value as the sixth similarity value. Alternatively, the electronic device can also take a smaller similarity value between the fourth similarity value and the fifth similarity value as the sixth similarity value.

[0291] Through any one of the methods for calculating the sixth similarity value, the electronic device can obtain the sixth similarity value of each input image for the target prompt word determined by combining the similarity value of the global image and the similarity value of the subject image.

[0292] For example, when the target prompt word is a prompt word including "normal exposure", the fourth similarity value of the input image numbered 1 is 0.2655, and the fifth similarity value of the subject object numbered 1' corresponding thereto is 0.2457. For example, by using weighted summation, the weight of the input image can be set to 0.6, and the weight corresponding to the subject image is set to 0.4. Then, the sixth similarity value of the input image is 0.2655*0.6+0.2457*0.4=0.2575. That is, the sixth similarity value of the input image numbered 1 for the target prompt word is 0.2575.

[0293] For example, the fourth similarity value of the input image numbered 4 is 0.8457, the fifth similarity value of the subject object numbered 4' corresponding thereto is 0.8874, for example, the weight of the input image can be set to 0.6 and the weight corresponding to the subject image is set to 0.4 by using weighted summation, then the sixth similarity value of the input image is 0.8457*0.6+0.8874*0.4=0.8623. That is, the sixth similarity value of the input image numbered 4 and the target prompt word is 0.8623.

[0294] For example, the fourth similarity value of the input image numbered 5 is 0.5412, the fifth similarity value of the subject object numbered 5' corresponding thereto is 0.4474, for example, the weight of the input image can be set to 0.6 and the weight corresponding to the subject image is set to 0.4 by using weighted summation, then the sixth similarity value of the input image is 0.5412*0.6+0.4474*0.4=0.5036. That is, the sixth similarity value of the input image numbered 5 and the target prompt word is 0.5036.

[0295] In this way, the electronic device can obtain the sixth similarity value of each input image and the target prompt word.

[0296] In S506, the electronic device determines the target image based on the sixth similarity value of each input image based on the preset threshold strategy.

[0297] In some actual demand scenarios, the target image refers to a normally exposed image.

[0298] In the scenario where the target prompt word is a prompt word containing “normal exposure”, the target image is the image with the highest similarity to the prompt word. The electronic device can obtain the image with the largest sixth similarity value as the target image. As in S505, the sixth similarity value of the input image numbered 4 is the largest, then the input image numbered 4 is the target image.

[0299] In the scenario where the target prompt word is a prompt word containing “underexposure” or the target prompt word is a prompt word containing “overexposure”, the target image that meets the normal exposure needs to be further obtained according to the threshold strategy.

[0300] In this embodiment, reference can be made to Figure 8The broken line diagram of the similarity values between the different prompt words and the images is given. The intersection of the broken line of the underexposure similarity value and the broken line of the overexposure similarity value can be understood as the similarity value point corresponding to the normally exposed image. That is, when the prompt word is the prompt word containing "underexposure", if the sixth similarity value of the image is within the numerical range corresponding to the intersection point, the image can be used as the target image; when the prompt word is the prompt word containing "overexposure", if the sixth similarity value of the image is within the numerical range corresponding to the intersection point, the image can be used as the target image.

[0301] Specifically, in some embodiments, when the target prompt word is the prompt word representing normal exposure, the exposure quantization label of the input image with the sixth similarity value greater than or equal to the first threshold value is a normally exposed image. The exposure quantization label of the input image with the sixth similarity value less than the first threshold value is a non-normally exposed image.

[0302] When the target prompt word is the prompt word representing underexposure, the exposure quantization label of the input image with the sixth similarity value greater than or equal to the second threshold value is an underexposed image. The exposure quantization label of the input image with the sixth similarity value less than the second threshold value and greater than or equal to the third threshold value is a normally exposed image, and the exposure quantization label of the input image with the sixth similarity value less than the third threshold value is an overexposed image.

[0303] When the target prompt word is the prompt word representing overexposure, the exposure quantization label of the input image with the sixth similarity value greater than or equal to the fourth threshold value is an overexposed image. The exposure quantization label of the input image with the sixth similarity value less than the fourth threshold value and greater than or equal to the fifth threshold value is a normally exposed image, and the exposure quantization label of the input image with the sixth similarity value less than the fifth threshold value is an underexposed image.

[0304] It can be understood that the above application scenarios can also be applied to the image exposure quantification method of the image comparison model without subject detection, that is, based on the fourth similarity value of the input image to the target prompt word and the threshold strategy corresponding to the target prompt word, the exposure quantification of the input image is performed to determine the target image matching the target prompt word. For specific methods, reference can be made to the method of determining the target image matching the target prompt word based on the sixth similarity value and the threshold strategy corresponding to the target prompt word in the above embodiments, which will not be described here.

[0305] Specifically, for example, when the target prompt word is the prompt word representing normal exposure, the first threshold value can be 0.7.

[0306] The exposure quantization label of the input image with the sixth similarity value greater than or equal to the first threshold value is a normally exposed image (the target image matching the target prompt word). The exposure quantization label of the input image with the sixth similarity value less than the first threshold value is a non-normally exposed image.

[0307] AsFigure 13 As shown, if the sixth similarity value of the input image numbered 4 is 0.885, the similarity value of the input image is greater than the first threshold value 0.7, and the similarity values of the remaining input images are all less than the first threshold value 0.7, the exposure quantization label of the input image numbered 4 is a normally exposed image, and the input image numbered 4 is the target image matched with the target prompt word. The exposure quantization labels of the remaining input images are non-normally exposed images.

[0308] Specifically, for example, when the target prompt word is an underexposed prompt word, the second threshold value and the third threshold value can be based on Figure 8 As shown, the trend of the line chart is determined. For example, as shown in FIG. 6, the trend of the line chart is determined by the trend line of the sixth similarity values of the input images. Figure 8 As shown, the intersection of the underexposed line and the overexposed point means the closest to normal exposure, and the second threshold value and the third threshold value are taken based on the similarity value corresponding to the intersection, for example, the second threshold value can be 0.55, and the third threshold value can be 0.45.

[0309] In one case, the exposure quantization label of the input image with the sixth similarity value greater than or equal to the second threshold value is an underexposed image, that is, the target image matched with the target prompt word.

[0310] As shown in FIG. 6, the sixth similarity value of the input image numbered 1 is 0.987, the sixth similarity value of the input image numbered 2 is 0.854, and the sixth similarity value of the input image numbered 3 is 0.712. The sixth similarity values of the input images numbered 1-3 are all greater than the second threshold value 0.55. The exposure quantization labels of the input images numbered 1-3 are underexposed images, that is, the target images matched with the target prompt word. Figure 13 In one case, the exposure quantization label of the input image with the sixth similarity value less than the second threshold value 0.55 and greater than or equal to the third threshold value 0.45 is a normally exposed image.

[0311] As shown in FIG. 6, the sixth similarity value of the input image numbered 4 is 0.501, which is greater than the third threshold value and less than the second threshold value, so the exposure quantization label of the input image numbered 4 is a normally exposed image.

[0312] Figure 13 In one case, the exposure quantization label of the input image with the sixth similarity value less than the third threshold value is an overexposed image.

[0313] As shown in FIG. 6, the sixth similarity value of the input image numbered 5 is 0.337. The sixth similarity value of the input image numbered 3 is less than the third threshold value 0.45. The exposure quantization label of the input image numbered 3 is an overexposed image.

[0314] As shown in FIG. 6, the sixth similarity value of the input image numbered 5 is 0.337. The sixth similarity value of the input image numbered 3 is less than the third threshold value 0.45. The exposure quantization label of the input image numbered 3 is an overexposed image. Figure 13

[0315] ​​The second threshold value and the third threshold value can also form a threshold range, and the electronic device can determine the image quantitative label based on the threshold range. The specific means used in this embodiment is not limited.

[0316] Specifically, for example, when the target prompt word is a prompt word representing overexposure, the fourth threshold value and the fifth threshold value can be based on Figure 8 As shown in the trend determination of the line chart, for example, as Figure 8 As shown, the intersection of the underexposure line and the overexposure line point means the closest normal exposure, and the fourth threshold value and the fifth threshold value are taken based on the similarity value corresponding to the intersection, for example, the fourth threshold value can be 0.58, and the fifth threshold value can be 0.43. In some embodiments, the fourth threshold value can be the same as the second threshold value, and the fifth threshold value can also be the same as the third threshold value.

[0317] In one case, the exposure quantitative label of the input image with the sixth similarity value greater than or equal to the fourth threshold value is an overexposure image.

[0318] As shown in the trend determination of the line chart, for example, as Figure 13 The sixth similarity value of the input image numbered 5 is 0.885, which is greater than the fourth threshold value 0.58, so the exposure quantitative label of the input image numbered 5 is an overexposure image, that is, a target image matching the target prompt word.

[0319] In one case, the exposure quantitative label of the input image with the sixth similarity value less than the fourth threshold value 0.58 and greater than or equal to the fifth threshold value 0.43 is a normally exposed image.

[0320] As shown in the trend determination of the line chart, for example, as Figure 13 The sixth similarity value of the input image numbered 4 is 0.557, which is less than the fourth threshold value and greater than the fifth threshold value. So the exposure quantitative label of the input image numbered 4 is a normally exposed image.

[0321] In one case, the exposure quantitative label of the input image with the sixth similarity value less than the fifth threshold value is an underexposure image.

[0322] As shown in the trend determination of the line chart, for example, as Figure 13 The sixth similarity value of the input image numbered 1 is 0.112, the sixth similarity value of the input image numbered 2 is 0.213, and the sixth similarity value of the input image numbered 3 is 0.323. The sixth similarity values of the input images numbered 1-3 are all less than the fifth threshold value 0.43. The exposure quantitative labels of the input images numbered 1-3 are underexposure images.

[0323] The second threshold value and the third threshold value can also form a threshold range, and the electronic device can determine the image quantitative label based on the threshold range. The specific means used in this embodiment is not limited.

[0324] In this way, based on the difference in target prompt words and the difference in threshold strategies corresponding to the target prompt words, the target image can be determined based on the sixth similarity values ​​of the respective input images.

[0325] In this embodiment, the trained image-text comparison model can also be used to understand the semantic correlation between images and target prompts, enabling querying of target images based on prompts. For example, when the target image is a normally exposed image, the threshold strategy corresponding to the target prompt for normally exposed images in the method provided in the above embodiment can be applied. Based on the similarity values ​​between each input image and the target prompt, an image with a similarity value greater than normally exposed can be identified from multiple input images as the target image. Exposure quantification of the input image is achieved through the image-text comparison model, resulting in an objective and accurate quantification result that can meet image exposure quantification requirements in a variety of scenarios.

[0326] In order to better illustrate the actual effect that can be achieved by using the image exposure quantization method provided in the embodiment of the present application. Figure 14 Several manually screened illustrations showing underexposure are given. Figure 14 The several examples given are all images of special scenes, for example, images with a large area of ​​light background (shopping mall); or images with a large area of ​​black background.

[0327] In the exposure evaluation results of manual exposure evaluation, Figure 14 Images numbered 1 to 9 are all underexposed images.

[0328] Use images numbered 1 to 9 Figure 10 The provided image exposure quantization method performs exposure quantization judgment and determines the exposure status label of each image. The obtained results include that image numbered 7 is an overexposed image, and the remaining images are normally exposed images.

[0329] Use images numbered 1 to 9 Figure 12 The provided image exposure quantization method performs exposure quantization judgment to obtain the exposure status label of each image. The result is that image number 7 is overexposed; the remaining images are underexposed images.

[0330] Figure 14 Among the 9 images provided, the actual exposure status is that image number 7 is an overexposed image, and the remaining images are underexposed images.

[0331] Obviously, the exposure evaluation results obtained by manual exposure evaluation cannot accurately determine the overexposed images. Figure 10 The image exposure quantification method provided in the embodiment can at least identify the overexposed image No. 7. Figure 12The image exposure quantification method provided in the embodiments can effectively improve the accuracy of exposure quantification in special scenarios (for example, there is a large area of light background (such as a shopping mall) in the image; or, there is a large area of black background in the image). Figure 14 The image with a light background or a dark background is accurately quantified, and the exposure quantification result is that the image numbered 7 is an overexposed image, and the remaining images are underexposed images. Figure 12 The image exposure quantification method provided in the embodiments can effectively improve the accuracy of exposure quantification in special scenarios (for example, there is a large area of light background (such as a shopping mall) in the image; or, there is a large area of black background in the image).

[0332] Figure 15 A possible structural schematic diagram of an electronic device involved in the above embodiments is shown. Figure 15 The electronic device 1500 shown includes a processing module 1501 and a storage module 1502. In some other possible manners, the electronic device 1500 can also include an image acquisition module 1503 and a display module 1504.

[0333] The processing module 1501 can be a central processing unit (CPU), a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field programmable gate array (FPGA) or other programmable logic devices, transistor logic devices, hardware components or any combination thereof. The processor can include an application processor and a baseband processor. It can implement or execute various exemplary logical blocks, modules and circuits described in combination with the disclosure. The processor can also be a combination of computing functions, such as one or more microprocessor combinations, combinations of DSP and microprocessor, etc.

[0334] For example, the processing module 1501 can be a processor 110 as shown in Figure 6 The image acquisition module 1503 can be a camera 193 as shown in Figure 6 The display module 1504 can be a display screen 194 as shown in Figure 6 The storage module 1502 can be an internal memory 121 as shown in Figure 6 The electronic device 1500 provided in the embodiments of the present application can be an electronic device 100 as shown in Figure 6

[0335] The embodiments of the present application also provide a chip system (for example, a system on a chip (SoC)), such as a chip system 200 as shown in Figure 16 ​As shown, the chip system includes at least one processor 1601 and at least one interface circuit 1602. The processor 1601 and the interface circuit 1602 can be interconnected by a line. For example, the interface circuit 1602 can be used to receive a signal from another device (e.g., a memory of the electronic device). For another example, the interface circuit 1602 can be used to send a signal to another device (e.g., a camera of the electronic device or the processor 1601). Illustratively, the interface circuit 1602 can read an instruction stored in the memory and send the instruction to the processor 1601. When the instruction is executed by the processor 1601, the electronic device can perform various steps in the above-described embodiments. Of course, the chip system can also include other discrete devices, which are not limited in the embodiments of the present application.

[0336] The embodiments of the present application further provide a computer readable storage medium, which includes computer instructions, when the computer instructions are run on the above-described electronic device, the electronic device performs various functions or steps performed by the electronic device 100 in the above-described method embodiments.

[0337] The embodiments of the present application further provide a computer program product, when the computer program product is run on a computer, the computer performs various functions or steps performed by the electronic device 100 in the above-described method embodiments. For example, the computer can be the above-described electronic device 100.

[0338] Through the above description of the embodiments, those skilled in the art can clearly understand that, for the convenience and brevity of description, only the above-described division of the functional modules is taken as an example for illustration, and in actual application, the above-described functions can be completed by different functional modules according to needs, that is, the internal structure of the device is divided into different functional modules to complete all or part of the functions described above.

[0339] In several embodiments provided in the present application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the above-described device embodiments are only illustrative, for example, the division of the modules or units is only a logical function division, and actual implementation can have another division manner, for example, a plurality of units or components can be combined or integrated into another device, or some features can be ignored or not executed. In addition, the coupling or direct coupling or communication connection between the displayed or discussed ones can be through some interfaces, indirect coupling or communication connection between devices or units, which can be electrical, mechanical or other forms.

[0340] The units described as separate components may or may not be physically separate, and the components displayed as units may be a physical unit or multiple physical units, that is, may be located in one place, or also can be distributed to multiple different places. Part or all of the units can be selected according to actual needs to achieve the purpose of the embodiment scheme.

[0341] In addition, each functional unit in each embodiment of the present application can be integrated in one processing unit, or each unit can be physically present alone, or two or more units can be integrated in one unit. The integrated unit can be realized in the form of hardware or in the form of a software functional unit.

[0342] The integrated unit, if realized in the form of a software functional unit and sold or used as an independent product, can be stored in a readable storage medium. Based on such understanding, the technical scheme of the embodiments of the present application essentially or the part that contributes to the prior art or the whole or part of the technical scheme can be embodied in the form of a software product, which is stored in a storage medium and includes a plurality of instructions for causing an apparatus (which can be a single-chip microcomputer, a chip, etc.) or a processor to execute all or part of the steps of the method described in each embodiment of the present application. The foregoing storage medium includes: a U disk, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, and various program code storage media.

[0343] The above is only a specific embodiment of the present application, but the protection scope of the present application is not limited thereto, any change or replacement within the technical scope disclosed in the present application should be covered in the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.

Claims

1. A method of quantifying image exposure, characterized by, The method comprises the following steps: obtaining an input image; the input image comprises one or more images to be quantified for exposure; performing significant model inference on the input image to obtain a subject image corresponding to the input image; inputting the input image and the subject image corresponding to the input image into a text-image contrast model respectively to obtain semantic similarity values of the input image and a preset prompt word, and semantic similarity values of the subject image and the preset prompt word; the preset prompt word comprises one or more description fields representing exposure states; wherein the semantic similarity value of the input image corresponding to the preset prompt word can represent a first exposure state of the input image, and the semantic similarity value of the subject image corresponding to the preset prompt word can represent a second exposure state of the subject image; determining an exposure quantification label of the input image based on the semantic similarity value of the input image corresponding to the preset prompt word and the semantic similarity value of the subject image corresponding to the preset prompt word; wherein the determination of the exposure quantification label is related to the consistency of the first exposure state and the second exposure state; when the first exposure state and the second exposure state are inconsistent, the exposure quantification label of the input image is determined based on the semantic similarity value of the subject image corresponding to the preset prompt word; the exposure quantification label is used to represent that the input image is an underexposed image, a normally exposed image or an overexposed image.

2. The method of claim 1, wherein, The input image comprises one image to be quantified for exposure, and the preset prompt word comprises a plurality of prompt words describing different exposure states; The inputting the input image and the subject image corresponding to the input image into a text-image contrast model respectively to obtain semantic similarity values of the input image and a preset prompt word, and semantic similarity values of the subject image and the preset prompt word comprises: inputting the input image into the text-image contrast model to obtain a first similarity value of the input image corresponding to each of the preset prompt words, thereby obtaining a plurality of first similarity values; inputting the subject image into the text-image contrast model to obtain a second similarity value of the subject image corresponding to each of the preset prompt words; The determining an exposure quantification label of the input image based on the semantic similarity value of the input image corresponding to the preset prompt word and the semantic similarity value of the subject image and the preset prompt word comprises: determining a third similarity value of the input image based on the first similarity value and the second similarity value; determining the exposure quantification label of the input image based on the third similarity value.

3. The method of claim 2, wherein, The first similarity value comprises semantic similarity values of the input image corresponding to a plurality of preset prompt words, and the second similarity value comprises semantic similarity values of the subject image corresponding to a plurality of preset prompt words, The determining a third similarity value of the input image based on the first similarity value and the second similarity value comprises: For each of the preset prompt words, the first similarity value and the second similarity value are weighted and summed to obtain a third similarity value corresponding to each of the preset prompt words, and a plurality of third similarity values are obtained; The third similarity value is based on the third similarity value to determine the exposure quantization label of the input image, comprising: Based on the third similarity value corresponding to the exposure state of the preset prompt word, the exposure quantization label of the input image is determined.

4. The method of claim 2, wherein, The first similarity value includes the semantic similarity value of the input image corresponding to the plurality of preset prompt words, and the second similarity value includes the semantic similarity value of the subject image corresponding to the plurality of preset prompt words, The third similarity value is based on the first similarity value and the second similarity value, comprising: Obtain the first similarity value with the maximum value and the second similarity value with the maximum value; In the case where the preset prompt word corresponding to the maximum first similarity value is consistent with the preset prompt word corresponding to the maximum second similarity value, the maximum first similarity value and the maximum second similarity value are weighted and summed to obtain the third similarity value; The third similarity value is based on the third similarity value to determine the exposure quantization label of the input image, comprising: Based on the third similarity value corresponding to the exposure state of the preset prompt word, the exposure quantization label of the input image is determined.

5. The method of claim 4, wherein, The method further comprises: If the preset prompt word corresponding to the maximum first similarity value is inconsistent with the preset prompt word corresponding to the maximum second similarity value, the maximum second similarity value is taken as the third similarity value.

6. The method according to any one of claims 1-5, characterized in that, The preset prompt word includes a first prompt word, a second prompt word and a third prompt word, the first prompt word is a prompt word representing underexposure, the second prompt word is a prompt word representing normal exposure, and the third prompt word is a prompt word representing overexposure.

7. The method of claim 1, wherein, The input image includes a plurality of images to be quantized for exposure, and the preset prompt word includes a target prompt word of a target exposure state; The input image and the subject image corresponding to the input image are respectively input into the image-text comparison model to obtain the semantic similarity value of the input image and the preset prompt word, and the semantic similarity value of the subject image and the preset prompt word, comprising: Each of the input images is input into the image-text comparison model to obtain a fourth similarity value corresponding to the target prompt word of each of the input images; Each of the subject images is input into the image-text comparison model to obtain a fifth similarity value corresponding to the target prompt word of each of the subject images; The third similarity value is based on the semantic similarity value of the input image corresponding to the preset prompt word and the semantic similarity value of the subject image and the preset prompt word to determine the exposure quantization label of the input image, comprising: Based on the fourth similarity value corresponding to each of the input images and the fifth similarity value of the subject image corresponding to the input image, a sixth similarity value of the input image is determined; Determine an exposure quantization label of each of the input images based on a sixth similarity value of each of the input images and a preset threshold strategy.

8. The method of claim 7, wherein, The determination of the sixth similarity value of each of the input images based on the fourth similarity value corresponding to each of the input images and the fifth similarity value of the subject image corresponding to the input image comprises: Taking an average value of the fourth similarity value corresponding to each of the input images and the fifth similarity value of the subject image corresponding to the input image as the sixth similarity value of the input image; Or, Taking a weighted sum value of the fourth similarity value corresponding to each of the input images and the fifth similarity value of the subject image corresponding to the input image as the sixth similarity value of the input image.

9. The method according to claim 7 or 8, characterized in that, The preset threshold strategy is related to the target prompt word, The preset threshold strategy comprises: When the target prompt word represents normal exposure, the exposure quantization label of the input image with the sixth similarity value greater than or equal to a first threshold value is a normal exposure image, and the exposure quantization label of the input image with the sixth similarity value less than the first threshold value is a non-normal exposure image; When the target prompt word represents underexposure, the exposure quantization label of the input image with the sixth similarity value greater than or equal to a second threshold value is an underexposure image, the exposure quantization label of the input image with the sixth similarity value less than the second threshold value and greater than or equal to a third threshold value is a normal exposure image, and the exposure quantization label of the input image with the sixth similarity value less than the third threshold value is an overexposure image; When the target prompt word represents overexposure, the exposure quantization label of the input image with the sixth similarity value greater than or equal to a fourth threshold value is an overexposure image, the exposure quantization label of the input image with the sixth similarity value less than the fourth threshold value and greater than or equal to a fifth threshold value is a normal exposure image, and the exposure quantization label of the input image with the sixth similarity value less than the fifth threshold value is an underexposure image.

10. A method for training a graph-text model for image exposure quantification, the method comprising: receiving a plurality of images; and training the graph-text model using the plurality of images. The method further comprises: Obtaining a plurality of sample images; Performing saliency model inference on the plurality of sample images to obtain a sample subject image corresponding to each of the sample images; Inputting the plurality of sample images into an initial image-text contrast model to obtain a first training similarity value of each of the sample images and a plurality of first training prompt words; the plurality of first training prompt words comprise prompt words representing different exposure states; Inputting the sample subject images into the initial image-text contrast model to obtain a second training similarity value of each of the sample subject images and a plurality of second training prompt words; the plurality of second training prompt words comprise prompt words representing different exposure states; For each sample image, obtaining a first training similarity value with the maximum value and a second training similarity value with the maximum value of the sample subject image corresponding to the sample image; If the first training prompt corresponding to the first training similarity value with the maximum value of the sample image is consistent with the second training prompt corresponding to the second training similarity value with the maximum value in representing the exposure state, a third training similarity value is determined based on the first training similarity value with the maximum value and the second training similarity value with the maximum value; If the preset prompt corresponding to the first similarity value with the maximum value is inconsistent with the preset prompt corresponding to the second similarity value with the maximum value in representing the exposure state, the second training similarity value with the maximum value is taken as the third training similarity value; A training label of the sample image is determined based on the first training prompt or the second training prompt corresponding to the third training similarity value; A loss is calculated based on the training label of the sample image and a reference label of the sample image; the reference label is an actual exposure state of the sample image, and the reference label is used to represent that the sample image is an underexposed image, a normally exposed image or an overexposed image; The initial image-text contrast model is iteratively trained based on the loss until the initial image-text contrast model meets a training condition, and the image-text contrast model is obtained.

11. An electronic device comprising a memory, a processor, and a computer program stored on the memory, wherein the computer program, when executed by the processor, is arranged to perform the method of any one of claims 1 to 10. The processor executes the computer program to implement the method of any one of claims 1-9, and / or implement the method of claim 10.

12. A computer readable storage medium having stored thereon a computer program, characterized in that, The computer program is executed by the processor to implement the method of any one of claims 1-9, and / or implement the method of claim 10.

13. A computer program product comprising a computer program, characterized in that, The computer program is executed by the processor to implement the method of any one of claims 1-9, and / or implement the method of claim 10. The computer program is executed by the processor to implement the method of any one of claims 1-9, and / or implement the method of claim 10.