Video dynamic range enhancement method and device, and terminal equipment

By performing multiple exposure and normal exposure processing at intervals and using corresponding fusion models to process video frames, the problem of high resource consumption in the prior art is solved, and real-time video dynamic range enhancement and stability improvement are achieved.

CN120224025APending Publication Date: 2025-06-27RDA MICROELECTRONICS BEIJING
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510535074.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-25
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

When the existing video dynamic range enhancement method meets real-time requirements, it requires high image sensor performance and equipment computing power, resulting in large resource consumption.

Method used

By performing multi-exposure processing and normal exposure processing at intervals, the multi-exposure fusion model and the timing multi-frame fusion model are used to process multiple frames of different exposure images and high dynamic range images of the current frame and the previous N frames respectively, reducing the calculation amount and the number of exposed frames.

Benefits of technology

The requirements for image sensors and equipment computing power are reduced, and the effect of improving the dynamic range of videos is achieved in real time. At the same time, the video after the dynamic range is enhanced is more stable.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120224025A_ABST
    Figure CN120224025A_ABST
Patent Text Reader

Abstract

The invention discloses a video dynamic range enhancement method and device and terminal equipment, and the method comprises the steps: carrying out the multi-exposure processing and normal exposure processing at intervals when a video frame is collected; for the video frame obtained by the multi-exposure processing, performing fusion processing on a plurality of frames of different exposure images corresponding to the video frame by using a multi-exposure fusion model to obtain a high dynamic range image, and adding the high dynamic range image into a video frame sequence; and for the video frame obtained by the normal exposure processing, performing fusion processing on the corresponding current frame of normal exposure image and the corresponding first N frames of high dynamic range images in the video frame sequence by using a time sequence multi-frame fusion model to obtain a high dynamic range image, and adding the high dynamic range image into the video frame sequence # imgabs0 #. By using the scheme of the invention, the video dynamic range can be effectively improved, and the requirements of video dynamic range enhancement on the computing power of an image sensor and equipment are reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of image processing technology, and in particular to a video dynamic range enhancement method, an apparatus, and a terminal device. Background Art

[0002] Currently, many imaging devices can obtain multiple consecutive images of a target shooting scene with different exposures within a certain period of time. Fusing multiple consecutive images with different exposures, that is, video dynamic range enhancement, will significantly suppress the noise of the image, and can significantly enhance the details and dynamic range of the image, which is of great significance.

[0003] In existing video dynamic range enhancement methods, in order to meet the real-time requirement, an HDR (High Dynamic Range) image sensor is used to reduce the time consumption, and a device with higher computing power is used to meet the requirement of multi-exposure fusion for each frame. Summary of the Invention

[0004] Embodiments of this application provide a video dynamic range enhancement method, an apparatus, and a terminal device to reduce the requirements for image sensors and device computing power in video dynamic range enhancement.

[0005] On the one hand, embodiments of this application provide a video dynamic range enhancement method, and the method includes:

[0006] When collecting video frames, perform multi-exposure processing and normal exposure processing at intervals;

[0007] For the video frames obtained by multi-exposure processing, use a multi-exposure fusion model to fuse the corresponding multiple different-exposure images to obtain a high dynamic range image, and add the high dynamic range image to the video frame sequence;

[0008] For the video frames obtained by normal exposure processing, use a temporal multi-frame fusion model to fuse the corresponding current frame normal exposure image with the corresponding first N high dynamic range images in the video frame sequence to obtain a high dynamic range image, and add the high dynamic range image to the video frame sequence. .

[0009] Optionally, the performing multi-exposure processing and normal exposure processing at intervals includes:

[0010] Determine whether the current frame needs to be subjected to multi-exposure processing;

[0011] If it is needed, perform multi-exposure processing to obtain the corresponding multiple different-exposure images;

[0012] If it is not needed, perform normal exposure processing to obtain the corresponding 1 normal exposure image.

[0013] Optionally, determining whether the current frame requires multi-exposure includes: determining whether the current frame requires multi-exposure processing according to a set step size, where the step size is used to indicate the number of frames of normal exposure processing at intervals of multi-exposure processing.

[0014] Optionally, determining whether the current frame requires multi-exposure includes:

[0015] Determining the difference amount between the current frame and the previous frame;

[0016] If the difference amount is greater than a set threshold, it is determined that the current frame requires multi-exposure processing; otherwise, it is determined that the current frame does not require multi-exposure processing.

[0017] Optionally, determining the difference amount between the current frame and the previous frame includes: taking the difference between each pixel in the normal exposure image of the current frame and the corresponding pixel in the normal exposure image of the previous frame, and taking the mean of the absolute values of the differences as the difference amount between the current frame and the previous frame.

[0018] Optionally, the method further includes:

[0019] Storing the video frame sequence in a cache;

[0020] Obtaining the high-dynamic range images of the first N frames in the video frame sequence from the cache.

[0021] Optionally, using the multi-exposure fusion model to perform fusion processing on corresponding multiple frames of differently exposed images to obtain a high-dynamic range image includes: inputting the corresponding multiple frames of differently exposed images into the multi-exposure fusion model, and obtaining a high-dynamic range image according to the output of the multi-exposure fusion model.

[0022] Optionally, using the multi-exposure fusion model to perform fusion processing on corresponding multiple frames of differently exposed images to obtain a high-dynamic range image includes: inputting the corresponding multiple frames of differently exposed images and the high-dynamic range image of the previous frame corresponding to the video frame sequence into the multi-exposure fusion model, and obtaining a high-dynamic range image according to the output of the multi-exposure fusion model.

[0023] Optionally, the multi-exposure fusion model performs fusion processing on multiple frames of differently exposed images in an alignment fusion manner; or the multi-exposure fusion model is trained using a machine learning method.

[0024] Optionally, the temporal multi-frame fusion model performs fusion processing on the normal exposure image of the current frame and the high-dynamic range images of the first N frames corresponding to the video frame sequence in an alignment fusion manner; or the temporal multi-frame fusion model is trained using a machine learning method.

[0025] On the other hand, an embodiment of the present application provides a video dynamic range enhancement device, and the device includes:

[0026] An image acquisition module, configured to perform multi-exposure processing and normal exposure processing at intervals when acquiring video frames;

[0027] A fusion processing module, configured to, for the video frames obtained by multi-exposure processing, use a multi-exposure fusion model to fuse corresponding multiple frames of differently-exposed images to obtain a high-dynamic range image, and add the high-dynamic range image to the video frame sequence; for the video frames obtained by normal exposure processing, use a temporal multi-frame fusion model to fuse the corresponding current-frame normal exposure image with the corresponding previous N frames of high-dynamic range images in the video frame sequence to obtain a high-dynamic range image, and add the high-dynamic range image to the video frame sequence. .

[0028] Optionally, the image acquisition module includes:

[0029] A determination unit, configured to determine whether multi-exposure processing needs to be performed on the current frame;

[0030] An exposure unit, configured to perform multi-exposure processing to obtain corresponding multiple frames of differently-exposed images when the determination unit determines that multi-exposure processing needs to be performed on the current frame; and perform normal exposure processing to obtain a corresponding 1-frame normal exposure image when the determination unit determines that multi-exposure processing does not need to be performed on the current frame.

[0031] Optionally, the device further includes: a cache module, configured to cache the video frame sequence.

[0032] Optionally, the device further includes: a multi-exposure fusion model training module, and / or a temporal multi-frame fusion model training module;

[0033] The multi-exposure fusion model training module is configured to train and obtain the multi-exposure fusion model;

[0034] The temporal multi-frame fusion model training module is configured to train and obtain the temporal multi-frame fusion model.

[0035] On the other hand, an embodiment of the present application further provides a terminal device, and the terminal device includes the video dynamic range enhancement device described above.

[0036] On the other hand, an embodiment of the present application further provides a computer-readable storage medium, and the computer-readable storage medium is a non-volatile storage medium or a non-transitory storage medium, on which a computer program is stored, and when the computer program is run by a processor, it executes the video dynamic range enhancement method.

[0037] On the other hand, an embodiment of the present application also provides a computer program product, including a computer program / instructions, which implement the steps of the video dynamic range enhancement method when executed by a processor.

[0038] For the video dynamic range enhancement method and device provided by the embodiments of the present application, when performing video dynamic range enhancement, only some video frames in the video are subjected to multi-frame exposure and fusion, and the remaining video frames perform fusion processing on the single-frame exposure result and the high-dynamic range images of the previous one or more frames. Because only multi-exposure fusion processing is performed on some frames of the video, compared with performing multi-exposure fusion on each frame of the video, the calculation amount and the number of exposure frames required are greatly reduced. Therefore, even with ordinary image sensors and devices with low computing power, the effect of real-time enhancing the video dynamic range can be achieved. In addition, in the solution of the present application, most video frames are obtained by fusing the previous one or more frames of the video with the current frame, which can not only achieve the purpose of enhancing the dynamic range, but also utilize potential temporal information, making the video after dynamic range enhancement more stable. Description of the Drawings

[0039] Figure 1 is a flowchart of a video dynamic range enhancement method provided by an embodiment of the present application;

[0040] Figure 2 is a schematic diagram of a video temporal exposure and fusion model used in an embodiment of the present application;

[0041] Figure 3 is a schematic diagram of a video temporal exposure and fusion model in an existing video dynamic range enhancement solution;

[0042] Figure 4 is a schematic structural diagram of a video dynamic range enhancement device provided by an embodiment of the present application;

[0043] Figure 5 is another schematic structural diagram of a video dynamic range enhancement device provided by an embodiment of the present application;

[0044] Figure 6 is a schematic hardware structure diagram of a terminal device provided by an embodiment of the present application. Detailed Embodiments

[0045] To make the above objects, features, and beneficial effects of the present application more obvious and understandable, the following detailed description of the specific embodiments of the present application is provided in conjunction with the drawings.

[0046] It should be noted that the terms used in the embodiments of the present application are only for the purpose of describing specific embodiments and are not intended to limit this specification. The singular forms of "a", "said" and "the" used in the embodiments of the present application and the appended claims are also intended to include plural forms, unless the context clearly indicates other meanings. In addition, the "plurality" appearing in the embodiments of the present application refers to two or more.

[0047] In the video collection, most of the scenes are the following two situations:

[0048] (1) If the camera is fixed, the background information in the consecutive frames collected is often repeated during the acquisition process;

[0049] (2) If the camera moves, the collected information generally changes gradually, resulting in a lot of information redundancy in the previous and next frames, and relatively few complete mutations.

[0050] It can be seen that in these scenes, the consecutive frames of the video have more repeated information.

[0051] For the above scenes, continuous multiple exposures can obtain as much information as possible about the brightness levels in the actual scene. For example, when underexposed, information about the brightly shining light tubes can be obtained, and when overexposed, information about the dark corners can be obtained. Much of the rich information obtained from these multiple exposures is repeated in the following frames of the video, including information about highlights and dark areas that require underexposure and overexposure to capture. If after multiple exposures are performed on a frame, the next frame of the video uses the result of its normal exposure and is fused with the high dynamic range image fused from the previous multiple exposures, the frame can obtain information about highlights and dark areas without multiple exposures, thus obtaining a high dynamic range image.

[0052] Based on the above analysis, the embodiment of the present application provides a video dynamic range enhancement method and device. When performing video dynamic range enhancement, only part of the video segments in the video are subjected to multi-frame exposure and fusion, and the remaining video segments are subjected to normal exposure processing, and the normal exposure image of the current frame obtained by the normal exposure processing is fused with the corresponding previous frame or multiple frames of high dynamic range images in the video frame sequence. For the above scenario, while meeting the requirements for real-time video dynamic range enhancement, the requirements for image sensor performance and device computing power can be greatly reduced.

[0053] like Figure 1 FIG. 1 is a flow chart of a method for enhancing the video dynamic range provided by an embodiment of the present application, comprising the following steps:

[0054] Step 101, when capturing video frames, multiple exposure processing and normal exposure processing are performed at intervals.

[0055] That is to say, when collecting video frames, multi-exposure processing is only performed on some of the video segments, while normal exposure processing is performed on other video segments.

[0056] Specifically, determine whether the current frame needs to be multi-exposure processed; if so, perform multi-exposure processing to obtain multiple different exposure images corresponding to this multi-exposure processing; if not, perform normal exposure processing to obtain 1 normal exposure image corresponding to this normal exposure processing.

[0057] The multi-exposure processing refers to not only obtaining a normal exposure image, but also obtaining images at other exposure levels, such as underexposed images and overexposed images; moreover, there can be multiple different degrees of underexposure and overexposure for the underexposed and overexposed images, and the embodiments of the present application do not limit this.

[0058] It should be noted that for the convenience of calculation, the exposure levels and the number of frames of the different exposure images obtained by each multi-exposure processing can be the same, or of course different, and the embodiments of the present application do not limit this. In addition, each multi-exposure processing can obtain 1 normal exposure image and 1 or multiple other exposure level images.

[0059] In specific implementation, a judgment rule for whether to perform multi-exposure processing can be preset.

[0060] For example, in some embodiments, it can be determined whether the current frame needs to be multi-exposure processed according to a set step size, and the step size is used to indicate the number of frames of normal exposure processing between multi-exposure processing intervals. That is to say, multi-exposure processing is performed every few frames of normal exposure processing.

[0061] For another example, in some other embodiments, it can be determined whether the current frame needs to be multi-exposure processed according to the difference amount between the current frame and the previous frame. Specifically, if the difference amount is greater than a set threshold, it is determined that the current frame needs to be multi-exposure processed; otherwise, it is determined that the current frame does not need to be multi-exposure processed. The threshold can be determined according to empirical simulation tests, and the embodiments of the present application do not limit this.

[0062] The calculation of the difference amount can be determined by comparing the pixels of the normal exposure image of the current frame with the corresponding normal exposure image of the previous frame in the video frame sequence. For example, in a non-limiting embodiment, the pixels of the normal exposure image of the current frame can be subtracted from the corresponding pixels of the normal exposure image of the previous frame, and the mean value of the absolute values of the differences is used as the difference amount between the current frame and the previous frame. Of course, other calculation methods can also be used to calculate the difference between the two, and the embodiments of the present application do not limit this.

[0063] The video frame sequence refers to the video frame sequence obtained after fusion processing.

[0064] Step 102: For the video frames obtained by multi-exposure processing, use the multi-exposure fusion model to fuse the corresponding multiple frames of differently exposed images to obtain a high-dynamic range image, and add the high-dynamic range image to the video frame sequence.

[0065] In some embodiments, the multi-exposure fusion model may adopt a traditional alignment and fusion method to fuse multiple frames of differently exposed images.

[0066] In some other embodiments, the multi-exposure fusion model may be pre-trained using machine learning methods. Specifically, multiple frames of differently exposed images for the same scene can be collected, and the collected images can be labeled. For example, corresponding labels can be generated through manual image retouching or professional software to generate sample images. To further enrich the diversity of the samples and enable the multi-exposure fusion model to better adapt to various different scenes, a training sample set can also be generated using sample images for multiple scenes. Use the training sample set to train the corresponding multi-exposure fusion model.

[0067] In some other embodiments, some existing multi-exposure fusion models can also be used, and this application does not make any limitations in this regard.

[0068] For multiple video images obtained by multi-exposure processing, when fusing them, only these multi-exposure images can be considered, that is, input the multiple frames of differently exposed images of this multi-exposure processing into the multi-exposure fusion model, and obtain the fused high-dynamic range image according to the output of the multi-exposure fusion model. Add the high-dynamic range image to the video frame sequence, that is, use the high-dynamic range image as the current frame of the video frame sequence; or the high-dynamic range image of the previous frame can also be considered at the same time. The high-dynamic range image of the previous frame refers to the high-dynamic range image of the previous frame relative to the current frame in the video frame sequence. For example, input the multiple frames of differently exposed images obtained by this multi-exposure processing and the high-dynamic range image of the previous frame into the multi-exposure fusion model, obtain the high-dynamic range image according to the output of the multi-exposure fusion model, and add the high-dynamic range image to the video frame sequence as the current frame of the video frame sequence.

[0069] It should be noted that the input parameters of the multi-exposure fusion model corresponding to different inputs will also be different during training, but the training methods are similar.

[0070] Step 103: For the video frames obtained by normal exposure processing, use the temporal multi-frame fusion model to fuse the corresponding current frame normal exposure image with the corresponding previous N high-dynamic range images in the video frame sequence to obtain a high-dynamic range image, and add the high-dynamic range image to the video frame sequence. 。

[0071] Taking 2 frames as an example, the input of the temporal multi-frame fusion model is the current frame's normal exposure image and the corresponding previous frame's high-dynamic range image in the video frame sequence; the output is the high-dynamic range picture after fusing these input images.

[0072] Similarly, in some embodiments, the temporal multi-frame fusion model can adopt a traditional alignment fusion method to fuse images with multiple different exposure parameters.

[0073] In some other embodiments, the temporal multi-frame fusion model can be pre-trained using machine learning methods. The training process of the temporal multi-frame fusion model is similar to that of the multi-exposure fusion model and will not be elaborated here.

[0074] In some other embodiments, some existing similar image frame fusion models can also be used, and the embodiments of this application do not make any limitations in this regard.

[0075] It should be noted that the same multi-exposure fusion model or different multi-exposure fusion models can be used for each frame undergoing multi-exposure processing. For example, when collecting video frames, multi-exposure processing is performed on some frames at a set step size, and the set step size = 2.

[0076] Correspondingly, at the 0th frame, multi-exposure processing is performed. When fusing the obtained multi-frame images with different exposures, the first method described above is adopted, that is, the corresponding multi-frame images with different exposures are input into the first multi-exposure fusion model to obtain the high-dynamic range image of the 0th frame of the video frame sequence.

[0077] At the 1st frame, normal exposure processing is performed, and the high-dynamic range image of the 1st frame of the video frame sequence is obtained through the temporal multi-frame fusion model.

[0078] Similarly, at the 2nd frame, normal exposure processing is performed, and the high-dynamic range image of the 2nd frame of the video frame sequence is obtained through the temporal multi-frame fusion model.

[0079] At the 3rd frame, multi-exposure processing is performed to obtain the multi-frame images with different exposures corresponding to this multi-exposure processing. Since the high-dynamic range image of the 2nd frame has been obtained, therefore, the second method described above is adopted, and the multi-frame images with different exposures obtained from this multi-exposure processing and the high-dynamic range image of the 2nd frame are input into the second multi-exposure fusion model to obtain the high-dynamic range image of the 3rd frame of the video frame sequence.

[0080] Among them, the input of the first multi-exposure fusion model is images with different exposure levels obtained for the same scene within a very short continuous time. For example, the input is underexposed images, normally exposed images, and overexposed images. The output of the first multi-exposure fusion model is a high-dynamic-range image obtained by fusing images with different exposure levels. The input of the second multi-exposure fusion model includes not only images with different exposure levels obtained for the same scene within a very short continuous time, but also the high-dynamic-range image of the previous frame.

[0081] It should be noted that since real-time fusion processing is performed on each frame of the acquired images, after the high-dynamic-range image of the current frame is obtained, the high-dynamic-range image can be cached, that is, the video frame sequence is cached. Correspondingly, in the above steps 102 and 103, when a high-dynamic-range image of one or more previous frames is required, the corresponding high-dynamic-range image can be obtained from the cache.

[0082] In addition, in specific implementation, when it is necessary to determine whether the current frame needs multi-exposure processing according to the difference between the current frame and the previous frame, the normally exposed image of the previous frame also needs to be cached and compared and judged after the normally exposed image of the current frame is obtained. Whether the judgment result is to perform multi-exposure processing or normal exposure processing, the normally exposed image of the current frame is used to update the cached normally exposed image of the previous frame for use in the exposure processing of the next frame. The following further gives an example to illustrate in detail the process of video image acquisition and processing using the video dynamic range enhancement method of the embodiments of the present application.

[0083] As Figure 2 shown, it is a schematic diagram of a video timing exposure and fusion model used in the embodiments of the present application.

[0084] In the 0th frame at the very beginning of the video, since there is no information of the previous frame, the 0th frame uses the multi-exposure method to obtain three images of underexposure, normal exposure, and overexposure, and uses the multi-exposure fusion model for fusion processing to obtain the fused high-dynamic-range image, which is output to the video frame sequence, that is, the 0th frame image of the video frame sequence is obtained.

[0085] At the 1st frame, since there is already the information of the 0th frame in the video frame sequence, the high-dynamic-range image of the 0th frame in the video frame sequence and the normally exposed image obtained at the 1st frame are input into the timing multi-frame fusion model for fusion processing to obtain the fused high-dynamic-range image, which is output to the video frame sequence, that is, the 1st frame image of the video frame sequence is obtained.

[0086] The image processing processes of the 2nd and 3rd frames are similar to that of the 1st frame.

[0087] At the 4th frame, three differently-exposed images obtained by multi-exposure processing can be input into the multi-exposure fusion model, or three differently-exposed images obtained by multi-exposure processing and the high-dynamic range image of the 3rd frame in the video frame sequence indicated by the dotted line can be used for fusion. If the high-dynamic range image of the 3rd frame in the video frame sequence is not added to the input of the multi-exposure fusion model, the multi-exposure fusion model used is the same as the multi-exposure fusion model used at the 0th frame, which is a model that fuses three images with different exposure levels. If the high-dynamic range image of the 3rd frame in the video frame sequence is added to the input of the multi-exposure fusion model, the multi-exposure fusion model used is different from the multi-exposure fusion model used at the 0th frame, and four images with different exposure levels need to be fused.

[0088] In the video dynamic range enhancement method provided by the embodiments of the present application, when performing video dynamic range enhancement, only some video segments in the video are subjected to multi-exposure processing, multiple multi-exposure images obtained by the multi-exposure processing are fused, and the remaining video segments are subjected to normal exposure processing. The single-frame normal exposure image obtained is fused with the high-dynamic range image of the previous frame or multiple frames corresponding in the video frame sequence. Because only some video segments are subjected to multi-exposure processing, compared with performing multi-exposure fusion on each frame of the video, the amount of calculation and the number of exposure output frames required are greatly reduced. Therefore, even with an ordinary image sensor and a device with low computing power, the effect of real-time enhancing the video dynamic range can be achieved.

[0089] Moreover, by using the video dynamic range enhancement method provided by the embodiments of the present application, not only can the amount of calculation be greatly reduced, but also the potential temporal information can be utilized, making the video after dynamic range enhancement more stable.

[0090] Take Figure 3 the schematic diagram of video temporal exposure and the fusion model used in the existing video dynamic range enhancement scheme shown as an example. Through comparison, the above effects of the solution of the present application can be better reflected.

[0091] Assume that the time taken for each exposure to output a frame is t. Take Figure 3 the first 6 frames of the video frame sequence shown as an example. In the existing video dynamic range enhancement scheme, each frame is subjected to multi-exposure (taking three different exposure levels of underexposure, normal exposure, and overexposure as an example), then Figure 3 the exposure output frame time for generating the 6 frames of the video frame sequence in the existing scheme shown is 18t.

[0092] According to the technical solution provided by the embodiments of the present application, as Figure 2 shown, the exposure output frame time for generating the 6 frames of the video frame sequence is 10t, which is only 56% of the time consumed by the existing technical scheme. Moreover, in the technical solution of the embodiments of the present application, by setting a larger multi-exposure step size, the time consumed for exposure to output frames can be further reduced.

[0093] In addition, from the perspective of the number of frames to be fused, in the existing video dynamic range enhancement solutions, 3 frames of images need to be fused each time. However, in the solution provided by the embodiments of the present application, in most cases, only 2 frames need to be fused, that is, the normal exposure image of the current frame and the high-dynamic range image of the previous frame are fused, which can further reduce the amount of calculation.

[0094] In addition, in the existing video dynamic range enhancement solutions, only the image information of different exposure levels of the current frame is considered during each fusion. However, in the solution proposed by the embodiments of the present application, when performing fusion, temporal information is also considered, and the high-dynamic range images of the previous frame or multiple frames are combined for fusion processing, which can make the resulting video images with enhanced dynamic range have better stability in terms of timing.

[0095] Correspondingly, the embodiments of the present application also provide a video dynamic range enhancement device, as Figure 4 shown, which is a schematic structural diagram of the video dynamic range enhancement device.

[0096] The video dynamic range enhancement device 400 of this embodiment includes: an image acquisition module 401 and a fusion processing module 402. Among them:

[0097] When the image acquisition module 401 is used to acquire video frames, multi-exposure processing and normal exposure processing are performed at intervals;

[0098] For the video frames obtained by multi-exposure processing, the fusion processing module 402 uses a multi-exposure fusion model to fuse the corresponding multi-frame differently exposed images to obtain a high-dynamic range image, and adds the high-dynamic range image to the video frame sequence; for the video frames obtained by normal exposure processing, the fusion processing module 402 uses a temporal multi-frame fusion model to fuse the corresponding normal exposure image of the current frame with the corresponding previous N high-dynamic range images in the video frame sequence to obtain a high-dynamic range image, and adds the high-dynamic range image to the video frame sequence. .

[0099] A non-limiting embodiment of the image acquisition module 401 may include: a judgment unit and an exposure unit. Among them:

[0100] The judgment unit is used to determine whether the current frame needs to be subjected to multi-exposure processing;

[0101] The exposure unit is used to perform multi-exposure processing to obtain corresponding multi-frame differently exposed images when the judgment unit determines that the current frame needs to be subjected to multi-exposure processing; and perform normal exposure processing to obtain a corresponding 1-frame normal exposure image when the judgment unit determines that the current frame does not need to be subjected to multi-exposure processing.

[0102] In some embodiments, the determination unit may determine whether the current frame needs to be multi-exposure processed according to a set step size, where the step size is used to indicate the number of frames of normal exposure processing at intervals of multi-exposure processing.

[0103] In other embodiments, the determination unit may determine whether the current frame needs to be multi-exposure processed according to the difference amount between the current frame and the previous frame. For the specific calculation and determination methods of the difference amount, reference may be made to the description in the method embodiments of the present application above, which will not be elaborated here.

[0104] As Figure 5 shown, in other embodiments, the video dynamic range enhancement device 400 may further include: a cache module 403 for caching the video frame sequence.

[0105] In some embodiments, the cache module 403 may also cache the previous frame of normal exposure image. Each time the current frame of normal exposure image is obtained through normal exposure, and after determining whether the current frame needs to be multi-exposure processed, the current frame of normal exposure image is updated to the cache, and the normal exposure image of this frame is used as the previous frame of normal exposure image during the next determination.

[0106] In some embodiments, the multi-exposure fusion model and the temporal multi-frame fusion model may perform fusion processing on the input multiple images in an alignment fusion manner.

[0107] In other embodiments, the multi-exposure fusion model and the temporal multi-frame fusion model may adopt existing fusion models with corresponding functions.

[0108] In other embodiments, the multi-exposure fusion model may also be trained in advance using machine learning methods. For example, the multi-exposure fusion model is trained using a multi-exposure fusion model training module, and the temporal multi-frame fusion model is trained using a temporal multi-frame fusion model training module. Correspondingly, the multi-exposure fusion model training module and / or the temporal multi-frame fusion model training module may be a part of the video dynamic range enhancement device in the embodiments of the present application, or may be independent of the video dynamic range enhancement device. The embodiments of the present application do not make any limitations in this regard.

[0109] Using the video dynamic range enhancement device provided by the embodiments of the present application, not only can the video dynamic range be enhanced, but also the amount of calculation and the number of exposure output frames required can be greatly reduced, and the requirements for the image sensor and the device computing power can be lowered. In addition, the video after dynamic range enhancement can be made more stable.

[0110] The video dynamic range enhancement method and device provided by the embodiments of the present application can be used on devices with a certain computing power, including but not limited to mobile phones, tablets, laptop computers, desktop computers, servers, etc.

[0111] Correspondingly, an embodiment of the present application further provides a terminal device, including the above video dynamic range enhancement device.

[0112] In specific implementation, for each module / unit included in each device and product described in the above embodiments, it may be a software module / unit, a hardware module / unit, or part may be a software module / unit and part may be a hardware module / unit.

[0113] For example, for each device and product applied to or integrated into a chip, each module / unit included therein may be implemented in a hardware manner such as a circuit, or at least part of the module / unit may be implemented in a software program manner, and the software program runs on a processor integrated inside the chip, and the remaining (if any) part of the module / unit may be implemented in a hardware manner such as a circuit; for each device and product applied to or integrated into a chip module, each module / unit included therein may be implemented in a hardware manner such as a circuit, and different modules / units may be located in the same component (such as a chip, a circuit module, etc.) or different components of the chip module, or at least part of the module / unit may be implemented in a software program manner, and the software program runs on a processor integrated inside the chip module, and the remaining (if any) part of the module / unit may be implemented in a hardware manner such as a circuit; for each device and product applied to or integrated into a terminal, each module / unit included therein may be implemented in a hardware manner such as a circuit, and different modules / units may be located in the same component (such as a chip, a circuit module, etc.) or different components inside the terminal, or at least part of the module / unit may be implemented in a software program manner, and the software program runs on a processor integrated inside the terminal, and the remaining (if any) part of the module / unit may be implemented in a hardware manner such as a circuit.

[0114] An embodiment of the present application also discloses a storage medium, which is a computer-readable storage medium, on which a computer program is stored, and when the computer program runs, it can execute the steps of the method shown in the foregoing embodiments. The storage medium may include a read-only memory (ROM), a random access memory (RAM), a magnetic disk, an optical disk, etc. The storage medium may also include a non-volatile memory or a non-transitory memory, etc.

[0115] An embodiment of the present application further provides a computer program product, including a memory and a processor, where a computer program that can run on the processor is stored on the memory, and when the processor runs the computer program, it executes the steps in the above method embodiments.

[0116] Please refer to Figure 6 , this embodiment of the present application also provides a schematic diagram of the hardware structure of a terminal device. The device includes a processor 601, a memory 602, and a transceiver 603.

[0117] The processor 601 may be a general-purpose central processing unit (CPU), a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits for controlling the execution of the program of the present application solution. The processor 601 may also include multiple CPUs, and the processor 601 may be a single-core (single-CPU) processor or a multi-core (multi-CPU) processor. Here, the processor may refer to one or more devices, circuits, or processing cores for processing data (such as computer program instructions).

[0118] The memory 602 may be a ROM or other types of static storage devices that can store static information and instructions, a RAM, or other types of dynamic storage devices that can store information and instructions, or may also be an electrically erasable programmable read-only memory (EEPROM), a compact disc read-only memory (CD-ROM), or other optical disc storage, optical disc storage (including compact discs, laser discs, optical discs, digital versatile discs, Blu-ray discs, etc.), magnetic storage media, or other magnetic storage devices, or any other medium that can be used to carry or store the desired program code in the form of instructions or data structures and can be accessed by a computer. This embodiment of the present application places no restrictions on this. The memory 602 may exist independently or may be integrated with the processor 601. Among them, the memory 602 may contain computer program code. The processor 601 is used to execute the computer program code stored in the memory 602, thereby implementing the method provided by this embodiment of the present application.

[0119] The processor 601, the memory 602, and the transceiver 603 are connected through a bus. The transceiver 603 is used to communicate with other devices or communication networks.

[0120] When Figure 6 the shown schematic diagram is used to illustrate the structure of the terminal device involved in the above embodiment, the processor 601 is used to control and manage the actions of the terminal device. For example, the processor 601 is used to support the terminal device to execute Figure 1Steps in, and / or actions performed by a terminal device in other processes described in embodiments of the present application. The processor 601 can communicate with other network entities through the transceiver 603. The memory 602 is used to store program codes and data of the terminal device.

[0121] It should be understood that the term "and / or" in this document is merely a description of the association relationship of associated objects, indicating that there can be three relationships. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone. In addition, the character " / " in this document indicates that the associated objects before and after are in an "or" relationship.

[0122] The term "a plurality of" that appears in embodiments of the present application refers to two or more.

[0123] The descriptions such as first and second that appear in embodiments of the present application are only for schematic and differentiating description objects, without an order, and do not represent a special limitation on the number of devices in embodiments of the present application, and cannot constitute any limitation on embodiments of the present application.

[0124] Each embodiment provided in the present application can be implemented in whole or in part by software, hardware, firmware, or any other combination. When implemented using software, the above embodiments can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions or computer programs. When the computer instructions or computer programs are loaded or executed on a computer, the processes or functions described in embodiments of the present application are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions can be transmitted from a website, computer, server, or data center to another website, computer, server, or data center in a wired or wireless manner. It should be understood that in various embodiments of the present application, the magnitudes of the serial numbers of the above processes do not mean the order of execution, and the order of execution of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of embodiments of the present application.

[0125] In several embodiments provided by the present application, it should be understood that the disclosed methods and apparatuses can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for example, the division of the units is only a logical function division, and there may be other division methods in actual implementation; for example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection to each other can be through some interfaces, and the indirect coupling or communication connection of the devices or units can be in electrical, mechanical or other forms.

[0126] In each embodiment of the present application, the functional units can be integrated in a processing unit, or each unit can be physically arranged separately, or two or more units can be integrated in one unit. The above integrated units can be implemented in the form of hardware, or in the form of a combination of hardware and software functional units.

[0127] Although the present application is disclosed as above, the present application is not limited thereto. Any person skilled in the art can make various changes and modifications without departing from the spirit and scope of the present application. Therefore, the protection scope of the present application should be subject to the scope defined by the claims.

Claims

1. A video dynamic range enhancement method, characterized in that: The method comprises: When capturing video frames, multiple exposure processing and normal exposure processing are performed at intervals; For the video frames obtained by multi-exposure processing, the corresponding multiple frames of different exposure images are fused using a multi-exposure fusion model to obtain a high dynamic range image, and the high dynamic range image is added to the video frame sequence; For the video frames obtained by normal exposure processing, the corresponding current frame normal exposure image is fused with the corresponding first N frames of high dynamic range images in the video frame sequence using a temporal multi-frame fusion model to obtain a high dynamic range image, and the high dynamic range image is added to the video frame sequence. .

2. The method according to claim 1, characterized in that The interval multi-exposure processing and normal exposure processing include: Determine whether the current frame needs to be processed by multiple exposures; If necessary, multi-exposure processing is performed to obtain corresponding multi-frame images with different exposures; If not necessary, normal exposure processing is performed to obtain a corresponding 1-frame normal exposure image.

3. The method according to claim 2, characterized in that Determining whether the current frame needs to be multi-exposure includes: Whether the current frame needs to be subjected to multi-exposure processing is determined according to a set step length, wherein the step length is used to indicate the number of frames of normal exposure processing in the multi-exposure processing interval.

4. The method according to claim 2, characterized in that: Determining whether the current frame needs to be multi-exposure includes: Determine the difference between the current frame and the previous frame; If the difference is greater than a set threshold, it is determined that the current frame needs to be processed by multiple exposures; otherwise, it is determined that the current frame does not need to be processed by multiple exposures.

5. The method according to claim 4, characterized in that Determining the difference between the current frame and the previous frame includes: The difference between each pixel in the current frame of normal exposure image and the corresponding pixel in the previous frame of normal exposure image is calculated, and the average of the absolute values ​​of the differences is taken as the difference between the current frame and the previous frame.

6. The method according to claim 1, characterized in that The method further comprises: Storing the video frame sequence in a buffer; The high dynamic range images of the first N frames in the video frame sequence are obtained from the cache.

7. The method according to claim 1, characterized in that The method of using a multi-exposure fusion model to fuse corresponding multiple frames of images with different exposures to obtain a high dynamic range image includes: The corresponding multiple frames of images with different exposures are input into the multi-exposure fusion model, and a high dynamic range image is obtained according to the output of the multi-exposure fusion model.

8. The method according to claim 1, characterized in that The method of using a multi-exposure fusion model to fuse corresponding multiple frames of images with different exposures to obtain a high dynamic range image includes: The corresponding multiple frames of different exposure images and the high dynamic range image of the corresponding previous frame in the video frame sequence are input into the multi-exposure fusion model, and the high dynamic range image is obtained according to the output of the multi-exposure fusion model.

9. The method according to any one of claims 1 to 8, characterized in that: The multi-exposure fusion model uses an alignment fusion method to fuse multiple frames of images with different exposures; or The multi-exposure fusion model is trained by using a machine learning method.

10. The method according to any one of claims 1 to 8, characterized in that: The time-series multi-frame fusion model uses an alignment fusion method to fuse the current frame normal exposure image with the corresponding first N frames of high dynamic range images in the video frame sequence; or The time series multi-frame fusion model is trained by using a machine learning method.

11. A video dynamic range enhancement device, characterized in that: The device comprises: An image acquisition module, used to perform multiple exposure processing and normal exposure processing at intervals when acquiring video frames; A fusion processing module is used to fuse the corresponding multiple frames of different exposure images using a multi-exposure fusion model for the video frames obtained by multi-exposure processing to obtain a high dynamic range image, and add the high dynamic range image to the video frame sequence; for the video frames obtained by normal exposure processing, a time-series multi-frame fusion model is used to fuse the corresponding current frame normal exposure image with the corresponding first N frames of high dynamic range images in the video frame sequence to obtain a high dynamic range image, and add the high dynamic range image to the video frame sequence. .

12. The device according to claim 11, characterized in that The image acquisition module comprises: A judgment unit, used to determine whether the current frame needs to be processed by multiple exposures; The exposure unit is used to perform multi-exposure processing when the judgment unit determines that the current frame needs to be processed by multi-exposure, so as to obtain corresponding multiple frames of differently exposed images; when the judgment unit determines that the current frame does not need to be processed by multi-exposure, perform normal exposure processing to obtain a corresponding one-frame normally exposed image.

13. The device according to claim 11, characterized in that The device also includes: A cache module is used to cache the video frame sequence.

14. The device according to any one of claims 11 to 13, characterized in that The device further comprises: a multi-exposure fusion model training module, and / or a temporal multi-frame fusion model training module; The multi-exposure fusion model training module is used to train and obtain the multi-exposure fusion model; The temporal multi-frame fusion model training module is used to train and obtain the temporal multi-frame fusion model.

15. A terminal device, characterized in that: The terminal device comprises the video dynamic range enhancement device as described in any one of claims 11 to 14.

16. A computer-readable storage medium, wherein the computer-readable storage medium is a non-volatile storage medium or a non-transient storage medium, and a computer program is stored thereon, wherein: When the computer program is executed by a processor, the steps of the video dynamic range enhancement method according to any one of claims 1 to 11 are executed.

17. A computer program product comprising a computer program / instructions, characterized in that When the computer program / instruction is executed by a processor, the steps of the video dynamic range enhancement method described in any one of claims 1 to 12 are implemented.