Image processing method and device, electronic equipment and storage medium

By semantic segmentation and weight acquisition of multiple frames of low dynamic range images and exposure fusion, the image problem of difficult to deal with brighter and darker areas in the prior art is solved, and high dynamic range images that are consistent with human eye perception are achieved efficiently.

CN120075639APending Publication Date: 2025-05-30BEIJING XIAOMI MOBILE SOFTWARE CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311605435.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-11-28
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

The prior art is difficult to effectively process images in brighter and darker areas in the same scene, resulting in improper exposure and unable to meet the brightness needs of the subjective perception of the human eye.

Method used

By semantic segmentation of multiple frames of low dynamic range images, the weights of each segmented area are obtained, and exposure fusion is performed based on these weights to generate high dynamic range images.

Benefits of technology

It realizes the generation of high dynamic range images that are more in line with the real perception of the human eye, with better results and precise control of the brightness of different regions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120075639A_ABST
    Figure CN120075639A_ABST
Patent Text Reader

Abstract

The invention relates to an image processing method and device, electronic equipment and a storage medium. The method comprises the following steps: acquiring multiple frames of first images, wherein the multiple frames of first images are low-dynamic-range images which are shot for the same scene and have different exposure degrees; semantic segmentation is carried out on the multiple frames of first images to obtain a segmentation result of each frame of first image, and the segmentation result comprises at least one segmentation area in the first image; obtaining a weight corresponding to each segmented region in each frame of the first image; and based on the weight corresponding to each segmented region in each frame of first image, performing exposure fusion on the plurality of frames of first images to obtain a second image which is a high dynamic range image. According to the method, when the weight is determined, the semantic features of the low-dynamic-range image are considered, so that the high-dynamic-range image obtained through fusion better conforms to real perception of human eyes, and a better effect is achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of image processing technologies, and in particular, to an image processing method, apparatus, electronic device, and storage medium. Background Art

[0002] When taking a photo, there are both brighter regions and darker regions in the same scene, and the brightness between the brighter regions and the darker regions may vary greatly. The dynamic range of the images that the image sensor of a camera can capture is usually limited. The longer the exposure time, the higher the image brightness, the dark areas are visible but the bright areas will be overexposed. The shorter the exposure time, the lower the image brightness, the dark areas are not visible but the details of the bright areas are visible. In order to prevent the bright areas of the image from being overexposed and the dark areas from being too dark, and the overall brightness conforms to the subjective perception of the human eye, usually multiple frames of low-dynamic-range images with different brightness levels are taken, and then a single frame of high-dynamic-range image is synthesized. Summary of the Invention

[0003] To overcome the problems existing in the related art, the present disclosure provides an image processing method, apparatus, electronic device, and storage medium.

[0004] According to a first aspect of an embodiment of the present disclosure, there is provided an image processing method, the method including:

[0005] Obtain multiple frames of first images, where the multiple frames of first images are low-dynamic-range images with different exposure degrees taken of the same scene;

[0006] Perform semantic segmentation on the multiple frames of first images to obtain the segmentation result of each frame of the first image, where the segmentation result includes at least one segmentation region in the first image;

[0007] Obtain the weight corresponding to each segmentation region in each frame of the first image;

[0008] Based on the weight corresponding to each segmentation region in each frame of the first image, perform exposure fusion on the multiple frames of first images to obtain a second image, where the second image is a high-dynamic-range image.

[0009] In some embodiments, the performing semantic segmentation on the multiple frames of first images to obtain the segmentation result of each frame of the first image includes:

[0010] Perform semantic segmentation on a target image to obtain the segmentation result of the target image, where the target image is an image in the multiple frames of first images that meets a preset condition;

[0011] Based on the positions of at least one segmentation region in the target image, determine the segmentation results of the other first images in the multiple frames of first images except the target image.

[0012] In some embodiments, the method further includes:

[0013] Obtaining the image brightness of each frame of the first image;

[0014] Sorting the multiple image brightnesses according to the brightness magnitude, and determining the image brightness at the middle position after sorting as the target image brightness;

[0015] Determining the first image corresponding to the target image brightness as the target image.

[0016] In some embodiments, the segmentation result further includes the category corresponding to each segmentation region; the obtaining the weight corresponding to each segmentation region in each frame of the first image includes:

[0017] Determining a reference image matching the multiple frames of the first image, and obtaining the reference region brightness of each reference segmentation region in the multiple reference segmentation regions of the reference image and the reference average brightness of the reference image;

[0018] Based on the target brightness, determining the weight corresponding to each segmentation region in each frame of the first image, where the target brightness includes the reference region brightness of each reference segmentation region or the reference average brightness of the reference image.

[0019] In some embodiments, the based on the target brightness, determining the weight corresponding to each segmentation region in each frame of the first image includes:

[0020] Based on the reference region brightness of the reference segmentation region, determining the weight corresponding to the segmentation region in each frame of the first image that belongs to the same category as the reference segmentation region;

[0021] When the category corresponding to at least one of the segmentation regions includes a target category and the category corresponding to the multiple reference segmentation regions does not include the target category, based on the reference average brightness, determining the weight corresponding to the segmentation region belonging to the target category in each frame of the first image.

[0022] In some embodiments, the determining a reference image matching the multiple frames of the first image includes:

[0023] Obtaining the image features of each frame of the first image;

[0024] Based on the image features of each frame of the first image and multiple reference image features, determining a target reference image feature in the multiple reference image features that matches the multiple image features;

[0025] Determining the image corresponding to the target reference image feature in the database as the reference image.

[0026] In some embodiments, obtaining the image features of each frame of the first image includes:

[0027] Performing feature extraction on each frame of the first image respectively to obtain the first features of each frame of the first image;

[0028] Based on the environmental information when each frame of the first image is captured, obtaining second features corresponding to the environmental information, where the environmental information is used to characterize the shooting environment when the first image is captured;

[0029] Based on the first features and the second features of each frame of the first image, determining the image features of each frame of the image.

[0030] In some embodiments, based on the target brightness, determining the weights corresponding to each segmentation region in each frame of the first image includes:

[0031] Obtaining the regional brightness of each segmentation region in each frame of the first image;

[0032] For each category, based on the target brightness corresponding to the category and the regional brightness of each segmentation region belonging to the category, determining a first segmentation region and a second segmentation region, where the target brightness is between the first regional brightness of the first segmentation region and the second regional brightness of the second segmentation region;

[0033] Based on the target brightness, the first regional brightness, and the second regional brightness, determining the weight corresponding to the first segmentation region and the weight corresponding to the second segmentation region;

[0034] Wherein, the weights corresponding to other segmentation regions belonging to the target category except the first segmentation region and the second segmentation region in the multiple frames of the first image are 0.

[0035] In some embodiments, after performing exposure fusion on the multiple frames of the first image based on the weights corresponding to each segmentation region in each frame of the first image to obtain a second image, the method further includes:

[0036] Based on the reference image, adjusting the brightness of the second image to obtain an adjusted second image.

[0037] In some embodiments, based on the reference image, adjusting the brightness of the second image to obtain an adjusted second image includes:

[0038] Based on the reference image, determining the segmentation regions to be adjusted in the second image;

[0039] Adjust the weight corresponding to the segmentation region in each first image that is in the same position as the segmentation region to be adjusted based on the region brightness of the segmentation region to be adjusted;

[0040] Perform exposure fusion on the multiple first images based on the weights corresponding to each segmentation region in the adjusted first images to obtain the adjusted second image.

[0041] According to a second aspect of the embodiments of the present disclosure, there is provided an image processing apparatus, the apparatus including:

[0042] An image acquisition module, configured to acquire multiple first images, the multiple first images being low dynamic range images with different exposure degrees taken of the same scene;

[0043] A semantic segmentation module, configured to perform semantic segmentation on the multiple first images to obtain the segmentation result of each first image, the segmentation result including at least one segmentation region in the first image;

[0044] A weight acquisition module, configured to acquire the weight corresponding to each segmentation region in each first image;

[0045] An image fusion module, configured to perform exposure fusion on the multiple first images based on the weights corresponding to each segmentation region in each first image to obtain a second image, the second image being a high dynamic range image.

[0046] In some embodiments, the semantic segmentation module is configured to:

[0047] Perform semantic segmentation on a target image to obtain the segmentation result of the target image, the target image being an image in the multiple first images that meets a preset condition;

[0048] Determine the segmentation results of the other first images in the multiple first images except the target image based on the positions of at least one segmentation region in the target image.

[0049] In some embodiments, the semantic segmentation module is configured to:

[0050] Acquire the image brightness of each first image;

[0051] Sort the multiple image brightnesses according to the brightness magnitude, and determine the image brightness at the middle position after sorting as the target image brightness;

[0052] Determine the first image corresponding to the target image brightness as the target image.

[0053] In some embodiments, the segmentation result further includes the category corresponding to each segmentation region; the weight acquisition module is configured to:

[0054] Determine a reference image that matches the multi-frame first images, and obtain the reference region brightness of each reference segmentation region in the multiple reference segmentation regions of the reference image and the reference average brightness of the reference image;

[0055] Based on the target brightness, determine the weight corresponding to each segmentation region in each frame of the first image, where the target brightness includes the reference region brightness of each reference segmentation region or the reference average brightness of the reference image.

[0056] In some embodiments, the weight acquisition module is configured to:

[0057] Based on the reference region brightness of the reference segmentation region, determine the weight corresponding to the segmentation region in each frame of the first image that belongs to the same category as the reference segmentation region;

[0058] When the category corresponding to the at least one segmentation region includes a target category and the category corresponding to the multiple reference segmentation regions does not include the target category, based on the reference average brightness, determine the weight corresponding to the segmentation region in each frame of the first image that belongs to the target category.

[0059] In some embodiments, the weight acquisition module is configured to:

[0060] Obtain the image feature of each frame of the first image;

[0061] Based on the image feature of each frame of the first image and multiple reference image features, determine the target reference image feature in the multiple reference image features that matches the multiple image features;

[0062] Determine the image corresponding to the target reference image feature in the database as the reference image.

[0063] In some embodiments, the weight acquisition module is configured to:

[0064] Extract features from each frame of the first image respectively to obtain the first feature of each frame of the first image;

[0065] Based on the environmental information when each frame of the first image is taken, obtain the second feature corresponding to the environmental information, where the environmental information is used to characterize the shooting environment when the first image is taken;

[0066] Based on the first feature and the second feature of each frame of the first image, determine the image feature of each frame of the image.

[0067] In some embodiments, the weight acquisition module is configured to:

[0068] Obtain the regional brightness of each segmentation region in each frame of the first image;

[0069] For each category, based on the target brightness corresponding to the category and the regional brightness of each segmentation region belonging to the category, determine a first segmentation region and a second segmentation region, where the target brightness is between the first regional brightness of the first segmentation region and the second regional brightness of the second segmentation region;

[0070] Based on the target brightness, the first regional brightness, and the second regional brightness, determine the weight corresponding to the first segmentation region and the weight corresponding to the second segmentation region;

[0071] Among them, the weights corresponding to other segmentation regions belonging to the target category in the multi-frame first image except the first segmentation region and the second segmentation region are 0.

[0072] In some embodiments, the device further includes:

[0073] A brightness adjustment module configured to adjust the brightness of the second image based on the reference image to obtain an adjusted second image.

[0074] In some embodiments, the brightness adjustment module is configured to:

[0075] Based on the reference image, determine the segmentation regions to be adjusted in the second image;

[0076] Based on the regional brightness of the segmentation regions to be adjusted, adjust the weights corresponding to the segmentation regions at the same positions as the segmentation regions to be adjusted in each frame of the first image;

[0077] Based on the weights corresponding to each segmentation region in each frame of the first image after adjustment, perform exposure fusion on the multi-frame first image to obtain the adjusted second image.

[0078] According to a third aspect of the embodiments of the present disclosure, there is provided an electronic device, including:

[0079] A processor;

[0080] A memory for storing processor-executable instructions;

[0081] Among them, the processor is configured to execute the method described in the first aspect of the embodiments of the present disclosure.

[0082] According to a fourth aspect of the embodiments of the present disclosure, a non-transitory computer-readable storage medium is provided. When the instructions in the storage medium are executed by a processor of an electronic device, the electronic device is enabled to execute the method described in the first aspect of the embodiments of the present disclosure.

[0083] Adopting the above method of the present disclosure has the following beneficial effects:

[0084] For the multi-frame low-dynamic range images, considering that the brightness of different regions in the same frame image that conforms to the subjective perception of the human eye is different, the method provided by the embodiments of the present disclosure first performs semantic segmentation on the multi-frame low-dynamic range images to obtain the segmentation results of each low-dynamic range image. The segmentation results include at least one segmentation region. Then, the weights corresponding to each segmentation region are respectively obtained. Then, based on the weights corresponding to each segmentation region in each low-dynamic range image, the multi-frame low-dynamic range images are subjected to exposure fusion to obtain a high-dynamic range image. This method considers the semantic features of the low-dynamic range images when determining the weights, so that the fused high-dynamic range image is more in line with the true perception of the human eye and has better effects.

[0085] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS

[0086] The accompanying drawings herein are incorporated into the specification and constitute a part of the specification, showing embodiments consistent with the present invention, and are used together with the specification to explain the principles of the present invention.

[0087] Figure 1 is a flowchart of an image processing method shown according to an exemplary embodiment;

[0088] Figure 2 is a flowchart of an image processing method shown according to an exemplary embodiment;

[0089] Figure 3 is a flowchart of an image processing method shown according to an exemplary embodiment;

[0090] Figure 4 is a block diagram of an image processing apparatus shown according to an exemplary embodiment;

[0091] Figure 5 is a block diagram of an electronic device shown according to an exemplary embodiment. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0092] Exemplary embodiments will be described in detail herein, and examples thereof are shown in the accompanying drawings. When the following description refers to the accompanying drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present invention. On the contrary, they are merely examples of devices and methods consistent with some aspects of the present invention as detailed in the appended claims.

[0093] In order to synthesize multiple captured low-dynamic-range (LDR) images with different brightness levels into a single high-dynamic-range (HDR) image. In a related art, based on the brightness, chromaticity, and local details of the LDR images, the weights corresponding to the LDR images are determined, and then, based on the weights corresponding to each frame of the LDR images, the multiple LDR images are subjected to exposure fusion to obtain the HDR image. In this method, only the local information of the LDR images is considered when determining the weights, resulting in the fused HDR image not conforming to the true perception of the human eye, and the effect of the HDR image being poor. Moreover, it is impossible to make a certain area brighter and another area darker. For example, when taking a portrait at night, it is desired that the face is brighter and the background is darker, but this exposure fusion method cannot achieve this effect.

[0094] In another related art, a deep learning approach is adopted. Multiple LDR images with different brightness levels are input into a deep learning network, and the HDR image is used as the groundtruth (manually annotated label). A deep learning network is trained through supervised learning. This deep learning network can output the HDR image based on the input multiple LDR images with different brightness levels. However, this deep learning network requires a large amount of training data, and the local brightness of the HDR image is uncontrollable. That is, when the deep learning network is trained, given a determined input image, the brightness of the output image is immutable. If the brightness of the output image needs to be adjusted, the groundtruth needs to be remade and the deep learning network needs to be retrained, which will change the brightness of all images. However, in practical applications, usually only the brightness of some images needs to be adjusted, while the brightness of other images remains unchanged. Therefore, the deep learning approach cannot achieve precise control of brightness.

[0095] In the embodiments of the present disclosure, in view of the problems in the related art above, when determining the weights, the semantic features of the low-dynamic-range images are considered, and the weights corresponding to each segmentation region in the low-dynamic-range images can be determined respectively, so that the fused high-dynamic-range image is more in line with the true perception of the human eye and has better effects. Moreover, compared with the deep learning method, this method in the present disclosure does not require a large amount of training data, and when the brightness needs to be adjusted, it can be achieved by adjusting the weights corresponding to the segmentation regions, and the brightness can be accurately controlled.

[0096] The method provided by the embodiments of the present disclosure is executed by an electronic device, which can be a device with image processing functions such as a mobile phone, a tablet computer, a notebook computer, a wearable device, a smart home device, a server, etc.

[0097] Figure 1 is a flowchart of an image processing method shown according to an exemplary embodiment, which is executed by an electronic device. Refer to Figure 1 , and the method includes the following steps:

[0098] Step S101, obtain multiple frames of first images, where the multiple frames of first images are low-dynamic-range images with different exposure degrees taken of the same scene.

[0099] Among them, the brightness of the multiple frames of first images is different, and the multiple frames of first images are taken of the same scene. There will also be differences in the environmental information when each frame of the first image is taken, and the environmental information may include information such as lighting conditions, weather, shooting time, shooting location, etc.

[0100] Step S102, perform semantic segmentation on the multiple frames of first images to obtain the segmentation results of each frame of the first image, where the segmentation results include at least one segmentation region in the first image.

[0101] Performing semantic segmentation on the first image means separating the objects of different categories in the first image to obtain the corresponding segmentation regions. Each segmentation region corresponds to one category, and at least one segmentation region can be segmented from one frame of the first image. Among them, the categories may include categories such as sky, green plants, human figures, animals, buildings, water, mountains, etc. For the multiple frames of first images, since the multiple frames of first images are taken of the same scene, the segmentation results of the multiple frames of first images are the same.

[0102] Step S103, obtain the weights corresponding to each segmentation region in each frame of the first image.

[0103] Since the objects contained in different segmented regions are different. For example, one segmented region contains a portrait, and another segmented region contains the sky. Different objects have different brightnesses that conform to the subjective perception of the human eye. To make the brightness of each object in the image conform to the subjective perception of the human eye, the weights corresponding to each segmented region in each frame of the first image are obtained respectively. The weight is used to characterize the importance of the corresponding segmented region during exposure fusion.

[0104] Step S104: Based on the weights corresponding to each segmented region in each frame of the first image, perform exposure fusion on multiple frames of the first image to obtain a second image, where the second image is a high-dynamic range image.

[0105] During exposure fusion, fusion is performed region by region, that is, according to the weights corresponding to each segmented region, the segmented regions at the same position in multiple frames of the first image are fused respectively to obtain the second image.

[0106] The method provided by the embodiments of the present disclosure is directed to multiple frames of low-dynamic range images. Considering that the brightnesses of different regions in the same frame of image conform to different subjective perceptions of the human eye, first perform semantic segmentation on multiple frames of low-dynamic range images to obtain the segmentation results of each frame of low-dynamic range image. The segmentation results include at least one segmented region. Then, obtain the weights corresponding to each segmented region respectively, and then based on the weights corresponding to each segmented region in each frame of low-dynamic range image, perform exposure fusion on multiple frames of low-dynamic range images to obtain a high-dynamic range image. In this way, when determining the weights, the semantic features of the low-dynamic range images are considered, so that the fused high-dynamic range image is more in line with the true perception of the human eye and has a better effect.

[0107] Figure 2 is a flowchart of an image processing method shown according to an exemplary embodiment, which is executed by an electronic device. Refer to Figure 2 and the method includes the following steps:

[0108] Step S201: Obtain multiple frames of the first image.

[0109] Among them, the brightnesses of multiple frames of the first image are different, and multiple frames of the first image are obtained by photographing the same scene. There will also be differences in the environmental information during the shooting of each frame of the first image. The environmental information may include information such as lighting conditions, weather, shooting time, and shooting location.

[0110] In some embodiments, when shooting the first image, record the environmental information during shooting.

[0111] Step S202: Perform semantic segmentation on multiple frames of the first image to obtain the segmentation results of each frame of the first image. The segmentation results include at least one segmented region in the first image.

[0112] Perform semantic segmentation on the first image, that is, segment objects of different categories in the first image to obtain corresponding segmentation regions, and each segmentation region corresponds to one category.

[0113] In some embodiments, the segmentation result includes the position information of each segmentation region, which is used to represent the position of the segmentation region in the first image; the segmentation result also includes the category corresponding to each segmentation region, and this category actually refers to the category of the objects included in the segmentation region.

[0114] In some embodiments, perform semantic segmentation on each frame of the first image to obtain the segmentation result of each frame of the first image.

[0115] In some embodiments, considering that some of the first images in multiple frames of the first image may be too bright or too dark, which may cause the objects in these first images to be unable to be accurately recognized when directly performing semantic segmentation. Therefore, semantic segmentation can be performed on the first images with medium brightness in multiple frames of the first images. Since the scenes captured by multiple frames of the first images are the same, the other first images can be segmented according to the segmentation result of the first image with medium brightness. That is, perform semantic segmentation on the target image to obtain the segmentation result of the target image, where the target image is the image in multiple frames of the first images that meets the preset conditions; based on the position of at least one segmentation region in the target image, determine the segmentation results of the other first images in multiple frames of the first images except the target image. Among them, the preset condition is that the image brightness of the images in multiple frames is medium brightness, and the medium brightness is the brightness that is neither too large nor too small among multiple image brightnesses.

[0116] Optionally, obtain the image brightness of each frame of the first image; sort the multiple image brightnesses according to the brightness magnitude, and determine the image brightness at the middle position after sorting as the target image brightness; determine the first image corresponding to the target image brightness as the target image. Among them, the middle position is the middle position in the sorted queue. When the number of multiple image brightnesses is odd and the middle position corresponds to one image brightness, this image brightness is used as the target image brightness. When the number of multiple image brightnesses is even and the middle position corresponds to two image brightnesses, any one of these two image brightnesses is used as the target image brightness.

[0117] Optionally, since the scenes captured by multiple frames of the first images are the same, the segmentation results of multiple frames of the first images should also be the same. Therefore, based on the position of at least one segmentation region in the target image, determine the segmentation results of the other first images in multiple frames of the first images except the target image. That is, based on the position of the segmentation region in the target image, segment the corresponding segmentation region at the same position in the other first images, and the segmentation regions at the same position in different first images belong to the same category.

[0118] In some embodiments, an image segmentation model is called to perform semantic segmentation on a first image (target image) to obtain a segmentation result. The image segmentation model can be a CNN (Convolutional Neural Network) or other types of network models, which are not limited in the embodiments of the present disclosure. Optionally, the image segmentation model outputs masks of multiple categories (i.e., segmentation regions) and category identifiers corresponding to each category.

[0119] Step S203: Determine a reference image that matches multiple frames of the first image, and obtain the reference region brightness of each reference segmentation region in the multiple reference segmentation regions of the reference image and the reference average brightness of the reference image.

[0120] The reference image is a pre-stored high-dynamic range image. The fact that the reference image matches multiple frames of the first image means that the categories of each segmentation region in the reference image and the brightness of the reference image are similar to those of the multiple frames of the first image.

[0121] In some embodiments, image features of each frame of the first image are obtained; based on the image features of each frame of the first image and multiple reference image features, a target reference image feature that matches the multiple image features among the multiple reference image features is determined; and the image corresponding to the target reference image feature in the database is determined as the reference image. The image features are used to describe the first image, the multiple reference image features are image features of pre-stored high-dynamic range images in the database, and the fact that the target reference image feature matches the multiple image features means that the similarity between the target reference image feature and the multiple image features is greater than a preset threshold.

[0122] In some embodiments, obtaining the image features of each frame of the first image includes: respectively performing feature extraction on each frame of the first image to obtain a first feature of each frame of the first image; based on the environmental information when each frame of the first image is captured, obtaining a second feature corresponding to the environmental information, where the environmental information is used to characterize the shooting environment when the first image is captured; and based on the first feature and the second feature of each frame of the first image, determining the image features of each frame of the image. The first feature is used to describe the image content of the first image, the second feature is used to describe the environmental information, and after the first feature and the second feature are combined, the obtained image features can more comprehensively describe the first image and are more accurate.

[0123] Optionally, a feature extraction network is called to perform feature extraction on the first image to obtain the first feature. The feature extraction network can use an image classification network as the backbone and use the output of the fully connected layer as the first feature.

[0124] Optionally, based on the first feature and the second feature of the first image in each frame, determine the image feature of each frame of image, including: splicing the first feature and the second feature to obtain the image feature.

[0125] It should be noted that, in some embodiments, the above method for obtaining the image feature can be used to obtain the reference image feature of the high-dynamic range image stored in the database, so as to construct the database.

[0126] In some embodiments, based on the image feature of the first image in each frame and multiple reference image features, determine the target reference image feature that matches the multiple image features among the multiple reference image features, including: based on the multiple image features, determine the target image feature, respectively determine the similarity between the target image feature and the multiple reference image features, and determine the reference image feature corresponding to the maximum similarity as the target reference image feature. Optionally, the target image feature is any one of the multiple image features; or, the target image feature is the image feature of the target image; or, the target image feature is the image feature after fusing the multiple image features.

[0127] It should be noted that in the embodiments of the present disclosure, only the case where the electronic device performs image retrieval to determine the reference image, the reference region brightness, and the reference average brightness is taken as an example. In some embodiments, when the amount of data such as the reference image and the reference image feature in the database is small, the electronic device can store the database, and then the electronic device performs image retrieval; when the amount of data such as the reference image and the reference image feature in the database is large, the database will occupy a large storage space. To save the storage space of the electronic device, the database can be stored in the cloud, and then the electronic device sends the image features of multiple frames of the first image to the cloud for image retrieval, and the cloud then sends the retrieval result to the electronic device.

[0128] Step S204, based on the target brightness, determine the weight corresponding to each segmentation region in each frame of the first image, where the target brightness includes the reference region brightness of each reference segmentation region or the reference average brightness of the reference image.

[0129] In some embodiments, determining the weight corresponding to each segmentation region in each frame of the first image based on the target brightness includes: obtaining the regional brightness of each segmentation region in each frame of the first image; for each category, based on the target brightness corresponding to the category and the regional brightness of each segmentation region belonging to the category, determining a first segmentation region and a second segmentation region, where the target brightness is between the first regional brightness of the first segmentation region and the second regional brightness of the second segmentation region, and the first regional brightness is the smallest one among at least one regional brightness greater than the target brightness, and the second regional brightness is the largest one among at least one regional brightness less than the target brightness; based on the target brightness, the first regional brightness, and the second regional brightness, determining the weight corresponding to the first segmentation region and the weight corresponding to the second segmentation region; wherein, the weights corresponding to other segmentation regions belonging to the target category except the first segmentation region and the second segmentation region in multiple frames of the first image are 0, and the sum of the weight corresponding to the first segmentation region and the weight corresponding to the second segmentation region is 1.

[0130] For example, the following formula is used to determine the weight corresponding to the first segmentation region and the weight corresponding to the second segmentation region:

[0131] W i =(T j -T) / (T j -T i )

[0132] W j =1 - W i

[0133] Wherein, W i is the weight corresponding to the second segmentation region, W j is the weight corresponding to the first segmentation region, T j is the first regional brightness, T i is the second regional brightness, and T is the target brightness.

[0134] In the above embodiments, the weights corresponding to each segmented region are determined. In some embodiments, for each pixel point in the segmented region, the weights corresponding to each pixel point can be determined respectively, that is, the pixel brightness of the pixel point is obtained; for each category, based on the target brightness corresponding to the category and the pixel brightness of the pixel point in the segmented region belonging to the category, a first pixel point and a second pixel point are determined, the target brightness is between the first pixel brightness of the first pixel point and the second pixel brightness of the second pixel point, and the first pixel point and the second pixel point are pixel points at the same position in multiple first images; based on the target brightness, the first pixel brightness, and the second pixel brightness, the weights corresponding to the first pixel point and the weights corresponding to the second pixel point are determined; wherein, the weights corresponding to other pixel points at the same position in the multiple first images except the first pixel point and the second pixel point are 0.

[0135] In some embodiments, considering that the categories of the reference segmented regions in the reference image may include the categories of all the segmented regions in the first image, or may not include the categories of all the segmented regions, for any segmented region in the first image, when the category of the segmented region is one of the categories of the multiple reference segmented regions, based on the reference region brightness of the reference segmented region, the weights corresponding to the segmented regions in each first image that belong to the same category as the reference segmented region are determined; when the category of the segmented region is not one of the categories of the multiple reference segmented regions, that is, when the category corresponding to at least one segmented region includes the target category and the categories corresponding to the multiple reference segmented regions do not include the target category, based on the reference average brightness, the weights corresponding to the segmented regions belonging to the target category in each first image are determined.

[0136] Step S205: Based on the weights corresponding to each segmented region in each first image, perform exposure fusion on the multiple first images to obtain a second image.

[0137] In the embodiments of the present disclosure, when performing exposure fusion, based on the multiple first images and the weights corresponding to each segmented region, a Laplacian pyramid corresponding to each region is constructed, and then, according to the weights corresponding to each segmented region, weighted summation is performed on the constructed Laplacian pyramid to obtain a fused Laplacian pyramid, and then the fused Laplacian pyramid is reconstructed to obtain a second image, that is, a high-dynamic-range image is obtained.

[0138] Step S206: Based on the reference image, adjust the brightness of the second image to obtain an adjusted second image.

[0139] After obtaining the second image, it is possible that the second image still does not fully conform to the subjective perception of the human eye, while the reference image used for reference is an image that conforms to the subjective perception of the human eye. Therefore, based on the reference image, the second image can be further fine-tuned to obtain an adjusted second image, making the second image more in line with the subjective perception of the human eye.

[0140] In some embodiments, based on the reference image, a segmentation region to be adjusted in the second image is determined. The segmentation region to be adjusted is a segmentation region in the second image where the reference region brightness of the reference segmentation region belonging to the same category as in the reference image differs significantly. Based on the region brightness of the segmentation region to be adjusted, the weights corresponding to the segmentation regions at the same position as the segmentation region to be adjusted in each frame of the first image are adjusted, while the weights corresponding to the other segmentation regions at different positions from the segmentation region to be adjusted in each first image do not need to be adjusted. Based on the weights corresponding to each segmentation region in each frame of the adjusted first image, exposure fusion is performed on multiple frames of the first image to obtain an adjusted second image.

[0141] Optionally, when the region brightness of the segmentation region to be adjusted is greater than the reference region brightness of the reference segmentation region belonging to the same category, the weight corresponding to the first segmentation region is increased, and the weight corresponding to the second segmentation region is decreased; when the region brightness of the segmentation region to be adjusted is less than the reference region brightness of the reference segmentation region belonging to the same category, the weight corresponding to the first segmentation region is decreased, and the weight corresponding to the second segmentation region is increased. Here, the first segmentation region and the second segmentation region are two regions corresponding to non-zero weights at the same position as the segmentation region to be adjusted in multiple frames of the first image.

[0142] It should be noted that in the embodiments of the present disclosure, only the brightness of the second image is adjusted based on the reference image as an example. In another embodiment, the user can observe the second image, and then according to the user's subjective perception, determine the region in the second image where the brightness needs to be adjusted, and adjust the weights corresponding to the corresponding segmentation regions in the first image.

[0143] In one example, Figure 2 The shown image processing method can also be as Figure 3 shown, see Figure 3, obtain multiple low-dynamic range images with different brightness levels; perform semantic segmentation on each low-dynamic range image respectively to obtain the segmentation result of each low-dynamic range image; perform image encoding on each low-dynamic range image respectively to obtain the image features of each low-dynamic range image; based on the image features of multiple low-dynamic range images, perform image retrieval in the database to detect the reference image and the reference region brightness of each reference segmentation region in the multiple reference segmentation regions of the reference image and the reference average brightness of the reference image; calculate the weight corresponding to each segmentation region in each frame of the first image; based on the calculated weights, perform exposure fusion on multiple low-dynamic range images to obtain a high-dynamic range image; then perform local fine-tuning on the high-dynamic range image, that is, adjust the weights corresponding to the segmentation regions to be adjusted; based on the adjusted weights, perform exposure fusion on multiple low-dynamic range images again to obtain an adjusted high-dynamic range image.

[0144] In the method provided by the embodiments of the present disclosure, when determining the weights, the semantic features of the low-dynamic range images are considered, so that the fused high-dynamic range image is more in line with the true perception of the human eye and has a better effect. Moreover, since each segmentation region is processed during processing, precise processing of different segmentation regions can be achieved, enabling each segmentation region in the fused high-dynamic range image to meet the true perception of the human eye and realizing precise brightness control.

[0145] Furthermore, after obtaining the high-dynamic range image, fine-tuning the high-dynamic range image can further improve the effect of the high-dynamic range image, making the adjusted high-dynamic range image more in line with the true perception of the human eye.

[0146] Figure 4 is a block diagram of an image processing apparatus shown according to an exemplary embodiment, configured in an electronic device, see Figure 4 , the apparatus includes:

[0147] An image acquisition module 401, configured to acquire multiple frames of first images, where the multiple frames of first images are low-dynamic range images with different exposure degrees taken of the same scene;

[0148] A semantic segmentation module 402, configured to perform semantic segmentation on multiple frames of first images to obtain the segmentation result of each frame of the first image, and the segmentation result includes at least one segmentation region in the first image;

[0149] A weight acquisition module 403, configured to acquire the weight corresponding to each segmentation region in each frame of the first image;

[0150] The image fusion module 404 is configured to perform exposure fusion on multiple frames of first images based on the weights corresponding to each segmentation region in each frame of the first image to obtain a second image, where the second image is a high-dynamic range image.

[0151] In some embodiments, the semantic segmentation module 402 is configured to:

[0152] Perform semantic segmentation on a target image to obtain a segmentation result of the target image, where the target image is an image in multiple frames of first images that meets a preset condition;

[0153] Based on the positions of at least one segmentation region in the target image, determine the segmentation results of the other first images in the multiple frames of first images except the target image.

[0154] In some embodiments, the semantic segmentation module 402 is configured to:

[0155] Obtain the image brightness of each frame of the first image;

[0156] Sort the multiple image brightnesses according to the brightness magnitude, and determine the target image brightness as the image brightness at the middle position after sorting;

[0157] Determine the first image corresponding to the target image brightness as the target image.

[0158] In some embodiments, the segmentation result further includes the category corresponding to each segmentation region; the weight acquisition module 403 is configured to:

[0159] Determine a reference image matching the multiple frames of first images, and obtain the reference region brightness of each reference segmentation region in the multiple reference segmentation regions of the reference image and the reference average brightness of the reference image;

[0160] Based on the target brightness, determine the weight corresponding to each segmentation region in each frame of the first image, where the target brightness includes the reference region brightness of each reference segmentation region or the reference average brightness of the reference image.

[0161] In some embodiments, the weight acquisition module 403 is configured to:

[0162] Based on the reference region brightness of the reference segmentation region, determine the weight corresponding to the segmentation region in each frame of the first image that belongs to the same category as the reference segmentation region;

[0163] When the category corresponding to at least one segmentation region includes a target category and the category corresponding to the multiple reference segmentation regions does not include the target category, based on the reference average brightness, determine the weight corresponding to the segmentation region in each frame of the first image that belongs to the target category.

[0164] In some embodiments, the weight acquisition module 403 is configured to:

[0165] Obtain the image features of the first image in each frame;

[0166] Based on the image features of the first image in each frame and multiple reference image features, determine the target reference image features in the multiple reference image features that match the multiple image features;

[0167] Determine the image corresponding to the target reference image feature in the database as the reference image.

[0168] In some embodiments, the weight acquisition module 403 is configured to:

[0169] Extract features from the first image in each frame respectively to obtain the first features of the first image in each frame;

[0170] Based on the environmental information when the first image in each frame is captured, obtain the second features corresponding to the environmental information, where the environmental information is used to characterize the shooting environment when the first image is captured;

[0171] Based on the first features and the second features of the first image in each frame, determine the image features of each frame of image.

[0172] In some embodiments, the weight acquisition module 403 is configured to:

[0173] Obtain the regional brightness of each segmentation region in the first image in each frame;

[0174] For each category, based on the target brightness corresponding to the category and the regional brightness of each segmentation region belonging to the category, determine the first segmentation region and the second segmentation region, where the target brightness is between the first regional brightness of the first segmentation region and the second regional brightness of the second segmentation region;

[0175] Based on the target brightness, the first regional brightness, and the second regional brightness, determine the weight corresponding to the first segmentation region and the weight corresponding to the second segmentation region;

[0176] Among them, the weights corresponding to other segmentation regions belonging to the target category except the first segmentation region and the second segmentation region in multiple frames of the first image are 0.

[0177] In some embodiments, the device further includes:

[0178] A brightness adjustment module, configured to adjust the brightness of the second image based on the reference image to obtain the adjusted second image.

[0179] In some embodiments, the brightness adjustment module is configured to:

[0180] Based on the reference image, determine the segmentation region to be adjusted in the second image;

[0181] Adjust the weight corresponding to the segmentation region at the same position as the segmentation region to be adjusted in each frame of the first image based on the regional brightness of the segmentation region to be adjusted;

[0182] Perform exposure fusion on multiple frames of the first image based on the weights corresponding to each segmentation region in the adjusted first image per frame to obtain an adjusted second image.

[0183] Regarding the device in the above embodiments, the specific manners in which each module performs operations have been described in detail in the embodiments related to the method, and will not be elaborated herein.

[0184] An embodiment of the present disclosure also provides an electronic device, including: a processor; a memory for storing processor-executable instructions; wherein the processor is configured to execute the image processing method in the above embodiments.

[0185] Figure 5 is a block diagram of an electronic device 500 shown according to an exemplary embodiment.

[0186] Refer to Figure 5 , the electronic device 500 may include one or more of the following components: a processing component 502, a memory 504, a power supply component 506, a multimedia component 508, an audio component 510, an input / output (I / O) interface 512, a sensor component 514, and a communication component 516.

[0187] The processing component 502 generally controls the overall operation of the electronic device 500, such as operations associated with display, telephone calls, data communication, camera operations, and recording operations. The processing component 502 may include one or more processors 520 to execute instructions to complete all or part of the steps of the above method. In addition, the processing component 502 may include one or more modules to facilitate the interaction between the processing component 502 and other components. For example, the processing component 502 may include a multimedia module to facilitate the interaction between the multimedia component 508 and the processing component 502.

[0188] The memory 504 is configured to store various types of data to support the operation of the electronic device 500. Examples of such data include instructions for any application or method operating on the electronic device 500, contact data, phone book data, messages, pictures, videos, etc. The memory 504 may be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, a magnetic disk, or an optical disk.

[0189] The power supply component 506 provides power for various components of the electronic device 500. The power supply component 506 may include a power management system, one or more power supplies, and other components associated with generating, managing, and distributing power for the electronic device 500.

[0190] The multimedia component 508 includes a screen that provides an output interface between the electronic device 500 and the user. In some embodiments, the screen may include a liquid crystal display (LCD) and a touch panel (TP). If the screen includes a touch panel, the screen can be implemented as a touch screen to receive input signals from the user. The touch panel includes one or more touch sensors to sense touches, swipes, and gestures on the touch panel. The touch sensors can not only sense the boundaries of the touch or swipe actions, but also detect the duration and pressure associated with the touch or swipe operation. In some embodiments, the multimedia component 508 includes a front camera and / or a rear camera. When the electronic device 500 is in an operating mode, such as a shooting mode or a video mode, the front camera and / or the rear camera can receive external multimedia data. Each of the front camera and the rear camera can be a fixed optical lens system or have a focal length and optical zoom capabilities.

[0191] The audio component 510 is configured to output and / or input audio signals. For example, the audio component 510 includes a microphone (MIC) that is configured to receive external audio signals when the electronic device 500 is in an operating mode, such as a call mode, a recording mode, and a voice recognition mode. The received audio signals can be further stored in the memory 504 or transmitted via the communication component 516. In some embodiments, the audio component 510 further includes a speaker for outputting audio signals.

[0192] The I / O interface 512 provides an interface between the processing component 502 and a peripheral interface module, and the peripheral interface module can be a keyboard, a click wheel, buttons, etc. These buttons may include, but are not limited to: a home button, a volume button, a power-on button, and a lock button.

[0193] The sensor assembly 514 includes one or more sensors for providing status assessments of various aspects of the electronic device 500. For example, the sensor assembly 514 can detect the on / off state of the electronic device 500, the relative positioning of components, such as the display and keypad of the electronic device 500. The sensor assembly 514 can also detect a change in the position of the electronic device 500 or a component of the electronic device 500, the presence or absence of user contact with the electronic device 500, the orientation or acceleration / deceleration of the electronic device 500, and a change in the temperature of the electronic device 500. The sensor assembly 514 can include a proximity sensor configured to detect the presence of nearby objects without any physical contact. The sensor assembly 514 can also include a light sensor, such as a CMOS or CCD image sensor, for use in imaging applications. In some embodiments, the sensor assembly 514 can also include an acceleration sensor, a gyroscope sensor, a magnetic sensor, a pressure sensor, or a temperature sensor.

[0194] The communication component 516 is configured to facilitate communication between the electronic device 500 and other devices in a wired or wireless manner. The electronic device 500 can access a wireless network based on communication standards, such as WiFi, 2G, or 3G, or a combination thereof. In an exemplary embodiment, the communication component 516 receives a broadcast signal or broadcast-related information from an external broadcast management system via a broadcast channel. In an exemplary embodiment, the communication component 516 further includes a near field communication (NFC) module to facilitate short-range communication. For example, the NFC module can be implemented based on radio frequency identification (RFID) technology, infrared data association (IrDA) technology, ultra-wideband (UWB) technology, Bluetooth (BT) technology, and other technologies.

[0195] In an exemplary embodiment, the electronic device 500 can be implemented by one or more application specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic components for performing the above-described method.

[0196] In an exemplary embodiment, a non-transitory computer-readable storage medium including instructions, such as the memory 504 including instructions, is also provided. The above instructions can be executed by the processor 520 of the electronic device 500 to complete the above-described method. For example, the non-transitory computer-readable storage medium can be a ROM, a random access memory (RAM), a CD-ROM, a magnetic tape, a floppy disk, and an optical data storage device, etc.

[0197] Embodiments of the present disclosure also provide a non-transitory computer-readable storage medium. When instructions in the storage medium are executed by a processor of an electronic device, the electronic device can execute the image processing method in the above embodiments.

[0198] Those skilled in the art will readily conceive of other embodiments of the present invention after considering the specification and practicing the invention disclosed herein. This application is intended to cover any variations, uses, or adaptations of the present invention, which follow the general principles of the present invention and include common general knowledge or conventional technical means in the technical field not disclosed in this disclosure. The specification and embodiments are only to be considered as exemplary, and the true scope and spirit of the present invention are pointed out by the following claims.

[0199] It should be understood that the present invention is not limited to the exact structures described above and shown in the drawings, and various modifications and changes can be made without departing from its scope. The scope of the present invention is only limited by the appended claims.

Claims

1. An image processing method, characterized in that, the method includes: obtaining multiple frames of first images, where the multiple frames of first images are low-dynamic range images with different exposure degrees taken of the same scene; performing semantic segmentation on the multiple frames of first images to obtain the segmentation result of each frame of the first image, where the segmentation result includes at least one segmentation region in the first image; obtaining the weight corresponding to each segmentation region in each frame of the first image; performing exposure fusion on the multiple frames of first images based on the weight corresponding to each segmentation region in each frame of the first image to obtain a second image, where the second image is a high-dynamic range image.

2. The method according to claim 1, characterized in that, performing semantic segmentation on the multiple frames of first images to obtain the segmentation result of each frame of the first image includes: performing semantic segmentation on a target image to obtain the segmentation result of the target image, where the target image is an image in the multiple frames of first images that meets a preset condition; determining the segmentation result of the other first images in the multiple frames of first images except the target image based on the position where at least one segmentation region in the target image is located.

3. The method according to claim 2, characterized in that, the method further includes: obtaining the image brightness of each frame of the first image; sorting multiple image brightnesses according to the brightness magnitude, and determining the image brightness at the middle position after sorting as the target image brightness; determining the first image corresponding to the target image brightness as the target image.

4. The method according to claim 1, characterized in that, the segmentation result further includes the category corresponding to each segmentation region; obtaining the weight corresponding to each segmentation region in each frame of the first image includes: determining a reference image matching the multiple frames of first images, and obtaining the reference region brightness of each reference segmentation region in multiple reference segmentation regions of the reference image and the reference average brightness of the reference image; determining the weight corresponding to each segmentation region in each frame of the first image based on a target brightness, where the target brightness includes the reference region brightness of each reference segmentation region or the reference average brightness of the reference image.

5. The method according to claim 4, characterized in that, determining the weight corresponding to each segmentation region in each frame of the first image based on a target brightness includes: determining the weight corresponding to the segmentation region in each frame of the first image that belongs to the same category as the reference segmentation region based on the reference region brightness of the reference segmentation region; when the category corresponding to the at least one segmentation region includes a target category and the category corresponding to the multiple reference segmentation regions does not include the target category, determining the weight corresponding to the segmentation region belonging to the target category in each frame of the first image based on the reference average brightness.

6. The method according to claim 4, characterized in that, determining a reference image matching the multiple frames of first images includes: obtaining the image feature of each frame of the first image; Determine target reference image features that match the multiple image features among the multiple reference image features based on the image features of each frame of the first image and the multiple reference image features; Determine the image corresponding to the target reference image feature in the database as the reference image.

7. The method according to claim 6, wherein, the obtaining the image features of each frame of the first image includes: Performing feature extraction on each frame of the first image respectively to obtain the first features of each frame of the first image; Based on the environmental information when each frame of the first image is captured, obtaining second features corresponding to the environmental information, where the environmental information is used to characterize the shooting environment when the first image is captured; Based on the first features and the second features of each frame of the first image, determining the image features of each frame of the image.

8. The method according to claim 4, wherein, the determining the weight corresponding to each segmentation region in each frame of the first image based on the target brightness includes: Obtaining the regional brightness of each segmentation region in each frame of the first image; For each category, based on the target brightness corresponding to the category and the regional brightness of each segmentation region belonging to the category, determining a first segmentation region and a second segmentation region, where the target brightness is between the first regional brightness of the first segmentation region and the second regional brightness of the second segmentation region; Based on the target brightness, the first regional brightness, and the second regional brightness, determining the weight corresponding to the first segmentation region and the weight corresponding to the second segmentation region; wherein, the weights corresponding to other segmentation regions belonging to the target category except the first segmentation region and the second segmentation region in the multiple frames of the first image are 0.

9. The method according to claim 4, wherein, after performing exposure fusion on the multiple frames of the first image based on the weight corresponding to each segmentation region in each frame of the first image to obtain a second image, the method further includes: Adjusting the brightness of the second image based on the reference image to obtain an adjusted second image.

10. The method according to claim 9, wherein, the adjusting the brightness of the second image based on the reference image to obtain an adjusted second image includes: Based on the reference image, determining the segmentation regions to be adjusted in the second image; Based on the regional brightness of the segmentation regions to be adjusted, adjusting the weights corresponding to the segmentation regions at the same positions as the segmentation regions to be adjusted in each frame of the first image; Based on the weights corresponding to each segmentation region in each frame of the first image after adjustment, performing exposure fusion on the multiple frames of the first image to obtain the adjusted second image.

11. An image processing apparatus, wherein, the apparatus includes: An image acquisition module configured to acquire multiple frames of first images, where the multiple frames of first images are low-dynamic range images with different exposure degrees of the same scene; A semantic segmentation module, configured to perform semantic segmentation on multiple frames of first images to obtain a segmentation result of each frame of the first images, where the segmentation result includes at least one segmentation region in the first images; A weight acquisition module, configured to acquire weights corresponding to each segmentation region in each frame of the first images; An image fusion module, configured to perform exposure fusion on the multiple frames of first images based on the weights corresponding to each segmentation region in each frame of the first images to obtain a second image, where the second image is a high-dynamic range image.

12. An electronic device, characterized in that, it includes: a processor; a memory for storing processor-executable instructions; wherein, the processor is configured to execute the method according to any one of claims 1-10.

13. A non-transitory computer-readable storage medium, characterized in that, when the instructions in the storage medium are executed by a processor of an electronic device, the electronic device is enabled to execute the method according to any one of claims 1-10.