A method and system for generating HDR video from multi-exposure image fusion
The HDR video generation system, which uses dual-camera acquisition and multi-exposure image fusion, solves the problems of frame rate improvement and image deblurring in low-light scenes, achieves high-quality generation of high dynamic range video, reduces hardware costs, and improves real-time performance.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-06
- Publication Date
- 2026-03-27
AI Technical Summary
Existing high dynamic range video generation methods do not adequately improve frame rate in low-light scenes, do not effectively solve the image deblurring problem, and have high hardware costs, making them difficult to popularize and resulting in unsatisfactory real-time performance.
A multi-exposure image fusion HDR video generation system was designed by using a dual-camera acquisition module to capture low, medium, and high exposure images, and then processing them with gamma correction and brightness alignment. The system utilizes an exposure time-guided attention module and a fusion network to generate HDR videos.
It enables high frame rate, high dynamic range video generation in low-light scenes, reduces artifacts, adapts to various scene requirements, lowers hardware costs, and improves video real-time performance and quality.
Smart Images

Figure CN119583964B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of video generation technology and relates to a method for generating high dynamic range video in low-light scenes using images with arbitrary exposure times as references. Specifically, it relates to a method and system for generating HDR video by fusing multiple exposure images. Background Technology
[0002] When generating high dynamic range (HDR) videos in low-light scenes, traditional single-exposure methods often fail to capture both shadow details and highlight areas simultaneously, limiting image quality and usability. To overcome this limitation, multi-exposure image fusion technology has gradually become a research hotspot. By capturing multiple images of the same scene at different exposure times, the dynamic range of the scene can be captured more comprehensively. However, multi-exposure image fusion in low-light environments still faces many challenges.
[0003] Traditional high dynamic range (HDR) video generation methods typically rely on specialized equipment that uses complex optical systems to distribute light across multiple sensors, setting different exposure parameters for each sensor to acquire multiple sets of low dynamic range (LDL) images. These images are then used to synthesize HDR video. While this method can capture a wider dynamic range, it requires expensive and complex hardware, making it difficult to popularize. Another common technique involves using a single camera to alternately capture images with different exposure times, then mapping these images to linear space and performing pixel-level motion alignment and fusion to generate HDR video frames. This method has lower hardware requirements and is suitable for more applications. However, in low-light scenes, long exposure times can cause moving objects in the image to blur or ghost, reducing image quality. Furthermore, this method is prone to ghosting during fusion and struggles to guarantee real-time video quality, especially with significant limitations in frame rate improvement.
[0004] With the development of deep learning technology, researchers have begun to leverage the powerful modeling capabilities of neural networks to address the challenges of high dynamic range (HDR) video generation. Neural networks can establish more accurate mapping relationships between multiple low dynamic range (LMR) images, enabling the calculation of optical flow fields between images and thus better aligning and integrating images with different exposures during the fusion process. These methods improve the quality of HDR videos while reducing hardware costs. However, existing HDR video generation methods still suffer from unsatisfactory real-time performance in low-light scenes. Most methods rely on medium-exposure images as references, which, while preserving more scene details, often requires longer exposure times in low-light conditions, leading to slower capture speeds, lower frame rates, and consequently affecting the real-time performance and overall visual quality of the video. Summary of the Invention
[0005] This invention addresses the problems of insufficient frame rate improvement and image deblurring in low-light scenes in existing high dynamic range video generation technologies. This invention provides a method and system for generating HDR video by fusing multiple exposure images. It can generate HDR video using images with arbitrary exposure times as references, can meet diverse scene requirements, and obtain high-quality high dynamic range video.
[0006] To achieve the above objectives, the following technical solution is adopted:
[0007] In one aspect, the present invention provides a method for generating HDR video by multi-exposure image fusion, comprising the following steps:
[0008] Step 1: Capture multi-exposure images using a dual-camera acquisition module.
[0009] The dual-camera acquisition module is designed to simultaneously capture low-exposure and medium-to-high-exposure image sequences to meet the requirements of high dynamic range video generation. The dual-camera acquisition module consists of camera 1 and camera 2, wherein:
[0010] Camera 1 (Low Exposure Camera): Used to capture high frame rate low exposure image streams to ensure that the system has a sufficient frame rate in low-light scenes to meet the requirements of video smoothness.
[0011] Camera 2 (Medium / High Exposure Camera): Alternates between capturing medium and high exposure images to complement video detail and dynamic range extension.
[0012] In low-light scenes, low-exposure image sequences are used as reference frames for generating HDR video. The dual-camera acquisition module transmits low, medium, and high-exposure image data acquired by camera 1 and camera 2 to the processing unit via a transmission interface (such as USB, WiFi, MIPI, etc.), ensuring efficient and stable data transmission.
[0013] Step 2: Frame synchronization and data preprocessing.
[0014] The processing unit is first responsible for synchronizing the acquired multi-exposure image data. Then, it performs data preprocessing on the synchronized multi-exposure images. The specific preprocessing steps include:
[0015] Gamma correction: Gamma correction is applied to low dynamic range images with different exposure times, mapping them to the high dynamic range domain to ensure uniformity of image brightness range.
[0016] Global brightness alignment: The brightness of images with different exposures is adjusted using a brightness alignment algorithm to make them consistent with the reference image.
[0017] After preprocessing, 3 pairs of 9-channel tensors (low, medium, and high exposures) are obtained. For low-light scenes, the tensors for low and medium exposures are swapped using a swapping strategy.
[0018] Step 3: Generate HDR video using an HDR video generation network based on multi-exposure image fusion.
[0019] The preprocessed tensor is used as input and fed into the HDR video generation network for multi-exposure image fusion to obtain reconstructed HDR frames. Multiple HDR frames are concatenated to obtain an HDR video.
[0020] Step 4: Video streaming;
[0021] The video stream generated by the HDR video generation network will be sent to the streaming module, which will then output the generated HDR video in different formats as needed.
[0022] The steps above will be explained in detail below:
[0023] Step 1 is as follows:
[0024] Camera 1 captures a continuous sequence of low-exposure images {L} with an exposure time of T1. l1 ,L l2 ,...,L ln Camera 2 captures medium and high exposure images alternately with exposure times T2 and T3 respectively. m ,L h}, using a low-exposure image subsequence as a reference for generating HDR video, based on the input subsequence {{L l1 ,L m ,L h},{L l2 ,L m ,L h},...,{L ln ,L m ,L h This will yield a continuous high dynamic range image subsequence output {I}. H1 ,I H2 ,...,I Hn The acquired image data consists of multiple sub-sequences. The dual-camera acquisition module transmits the low, medium, and high exposure image data to the processing unit via a transmission interface (such as USB, WiFi, MIPI, etc.), ensuring efficient and stable data transmission.
[0025] Step 2 is as follows:
[0026] Since the capture of low, medium, and high exposure images may not be synchronized, the processing unit first aligns the images transmitted by the dual-camera acquisition modules in time according to their timestamps. Then, data preprocessing is performed; for each frame of video output, three LDR images with different exposure times are required. l ,L m ,L hAs input, to facilitate alignment detection, the LDR image is first mapped to the HDR domain using gamma correction:
[0027]
[0028] Where γ is the gamma correction parameter, set to 2.2, t i Let G be the exposure time. Then, the gamma-corrected feature set {G} will be obtained. i}
[0029] LDR image with intermediate exposure time L m The non-reference image is used as a reference image. Subsequently, histogram equalization is applied to align the non-reference image with the reference image. Relative brightness value v i The definition is as follows:
[0030]
[0031] Where h and w represent the height and width of the image, respectively. This represents the brightness value at coordinates (x, y) in the LDR image. Global brightness alignment is achieved by mapping the LDR image to the brightness range of the reference image based on the relative brightness value, as expressed by the following formula:
[0032] A i =L i ·v i i = l, m, h;
[0033] Among them, A i This indicates the corresponding brightness-aligned output. It's worth noting that, due to L... m It is a reference image, A m =L m Finally, for each frame's three inputs, we can obtain 3 pairs of 9-channel tensors {{L}. l G l A l}、{L m G m A m}、{L h G h A h}}.
[0034] Then, a set of tensors with low-exposure images as references is obtained through a swapping strategy, as expressed by the following formula:
[0035]
[0036] in This represents a tensor that is swapped from low exposure to medium exposure using a swapping strategy. This represents a tensor that is swapped from medium exposure to low exposure using a swapping strategy.
[0037] According to this exchange strategy, the present invention can use any exposure time as a reference.
[0038] Furthermore, the HDR video generation network for multi-exposure image fusion consists of two parts: a fusion network and a restoration network. To utilize exposure time to promote the fusion of images with different exposure times, this invention constructs a fusion network based on a designed exposure time-guided attention module. For a set of LDR images {L... l ,L m ,L h}, and its corresponding exposure time is {t} low ,t mid ,t high After processing in step 2, a 9-channel tensor is obtained as the input to the fusion network. The input exposure time is standardized and relativized using the following formula:
[0039]
[0040] Among them, e low and e high represents the relative exposure time for low and high exposure inputs, respectively. c is a hyperparameter, set to 10. The exposure time-guided attention module employs a multi-scale structure, where the input features are downsampled three times sequentially, forming four different scales of input with the original scale input. The initial reference feature map and non-reference feature map at the j-th scale are respectively represented by _____. and The input feature maps at four different scales (j = 1, 2, 3, 4, where j = 1 represents the original scale) are processed by global average pooling, two fully connected layers, and a sigmoid activation function to generate reference and non-reference feature modulation coefficients. Meanwhile, the relative exposure times of the non-reference features are each passed through a fully connected layer to obtain the time modulation coefficient t. i ∈R 1×1×128 Then, the time modulation coefficient t i Combined with the corresponding non-reference feature modulation coefficients, and passed through three fully connected layers and a sigmoid activation function, the final non-reference feature modulation coefficients are obtained. The final output of the exposure time-guided attention module is:
[0041]
[0042] in, This serves as the output reference feature after the exposure time-guided attention module. The output non-reference features are those of the attention module guided by exposure time.
[0043] Existing attention fusion methods are used to fuse the features obtained above to obtain the output of the fusion network. Finally, an existing DomainPlus-based recovery network is used to further correct the fused features, resulting in high dynamic range (HDR) output frames. These HDR output frames are concatenated to obtain the reconstructed HDR video.
[0044] Furthermore, the network training steps in step 3 are as follows:
[0045] L1 loss is used as the basic loss function for pixel-level supervision, and the extended advanced Sobel loss function (D-ASL) is used as an additional function, which is defined as follows:
[0046]
[0047] Where ASL is the advanced Sobel loss, and N is the training batch size. Z represents the predicted output and ground truth value of the HDR video generation network for multi-exposure image fusion, respectively, and Sobel represents the advanced Sobel filter. The expansion rate is d i The ASL is calculated using the i = {1, 2, 3} setting. Following existing techniques, this invention employs the D-ASL to formulate the loss function. The final loss function is:
[0048]
[0049] Where λ is a hyperparameter. Optimization is performed using the Adam optimizer, with an initial learning rate set to 10. -4 When the learning rate is below 10 -6 At this point, training ends, and the weights of each layer of the network are continuously updated by the backpropagation algorithm.
[0050] On the other hand, the present invention provides an HDR video generation system for multi-exposure image fusion, including a dual-camera acquisition module, a processing unit, a video generation module, and a streaming module.
[0051] The dual-camera acquisition module is used to achieve multi-exposure image capture. The dual-camera acquisition module consists of camera 1 and camera 2, wherein:
[0052] Camera 1 (Low Exposure Camera): Used to capture high frame rate low exposure image streams to ensure that the system has a sufficient frame rate in low-light scenes to meet the requirements of video smoothness.
[0053] Camera 2 (Medium / High Exposure Camera): Alternates between capturing medium and high exposure images to complement video detail and dynamic range extension.
[0054] In low-light scenes, low-exposure image sequences are used as reference frames for generating HDR video. The dual-camera acquisition module transmits low, medium, and high-exposure image data acquired by camera 1 and camera 2 to the processing unit via a transmission interface (such as USB, WiFi, MIPI, etc.), ensuring efficient and stable data transmission.
[0055] The processing unit is used to perform frame synchronization and data preprocessing on the image data acquired by the dual-camera acquisition module.
[0056] The processing unit is first responsible for synchronizing the acquired multi-exposure image data. Since the capture of low, medium, and high-exposure images may not be synchronized, the system uses a synchronization mechanism to ensure that these images are aligned in time. Then, the synchronized multi-exposure images undergo data preprocessing, specifically including the following steps:
[0057] Gamma correction: Gamma correction is applied to low dynamic range images with different exposure times, mapping them to the high dynamic range domain to ensure uniformity of image brightness range.
[0058] Global brightness alignment: The brightness of images with different exposures is adjusted using a brightness alignment algorithm to make them consistent with the reference image.
[0059] After preprocessing, 3 pairs of 9-channel tensors (low, medium, and high exposures) are obtained. For low-light scenes, the tensors for low and medium exposures are swapped using a swapping strategy.
[0060] The video generation module uses a trained HDR video generation network based on multi-exposure image fusion to generate HDR videos.
[0061] The tensor preprocessed by the processing unit is used as input and fed into the HDR video generation network for multi-exposure image fusion to obtain reconstructed HDR frames. Multiple HDR frames are concatenated to obtain an HDR video.
[0062] The streaming module is responsible for outputting the HDR video generated by the video generation module in different formats.
[0063] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0064] This invention employs a strategy that uses images with arbitrary exposure times as a reference. By acquiring images with high frame rates and low exposures, as well as intermittent medium-to-high exposures, it achieves high dynamic range video generation and can flexibly adapt to various scenarios without being affected by exposure time on the generated video frame rate.
[0065] 1. By using preprocessing methods, we can mitigate the undesirable differences between LDR images captured at different exposure times in real-world scenes. We also propose a swapping strategy that allows us to select any image as a reference when generating high dynamic range videos, thus avoiding the impact of high exposure time on the frame rate of the generated video in low-light scenes.
[0066] 2. An exposure time-guided attention module was designed to promote the fusion of HDR images through exposure time, thereby further reducing artifacts in the generated high dynamic range video. Attached Figure Description
[0067] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments recorded in this invention. For those skilled in the art, other drawings can be obtained based on these drawings.
[0068] Figure 1 A schematic diagram of an HDR video generation system that uses multi-exposure image fusion;
[0069] Figure 2 This is a schematic diagram of the preprocessing unit in an embodiment of the present invention;
[0070] Figure 3 This is a schematic diagram of the exposure time-guided attention module structure in an embodiment of the present invention;
[0071] Figure 4 This is a diagram illustrating the architecture of the video generation system in an embodiment of the present invention.
[0072] Figure 5 , Figure 6 This is to demonstrate the effects of the present invention and its embodiments. Detailed Implementation
[0073] To better understand this technical solution, the technical solution of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described examples are only a part of the embodiments of the present invention, and not all of them. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art based on this application are within the scope of protection of the present invention.
[0074] like Figure 1 As shown, a method for generating HDR video by multi-exposure image fusion includes the following steps:
[0075] Step 1: Dual-camera acquisition module and multi-exposure image capture.
[0076] This system architecture includes a dual-camera acquisition module designed to simultaneously capture low-exposure and medium-to-high-exposure image sequences to meet the requirements of high dynamic range video generation. The dual-camera acquisition device consists of two standard industrial cameras, model MV-CS032-10GC. Although this example uses a dual-camera design, this design can be extended to other integrated solutions. One camera captures a continuous low-exposure image subsequence L with an exposure time of 30,000 microseconds. l1 ,L l2 ,…,L ln This ensures that the minimum FPS requirement is met. Simultaneously, a second camera alternately captures medium and high exposure images L with exposure times of 120,000 microseconds and 480,000 microseconds respectively. m ,L h Instead of continuous shooting, a sequence of low-exposure images is used as reference frames for generating HDR video, based on the input subsequence {{L}. 11 ,L m ,L h},{L 12 ,L m ,L h},…,{L 1n ,L m ,L h This will yield a continuous high dynamic range image subsequence output {I}. H1 ,I H2 ,...,I Hn The acquired image data consists of multiple subsequences, which are used to generate HDR videos.
[0077] Step 2: Frame synchronization and data preprocessing.
[0078] The acquired image sequences reach the processing unit via the Wi-Fi module and are synchronized using timestamps to ensure that the images transmitted by the dual-camera acquisition modules are time-aligned. Then, data preprocessing is performed, such as... Figure 2 As shown, for each frame of video output, three LDR images with different exposure times are required as input {L l ,L m ,L h To facilitate alignment detection, the LDR image is first mapped to the HDR domain using gamma correction.
[0079]
[0080] Where γ is the gamma correction parameter, set to 2.2, t i Let G be the exposure time. Then, the gamma-corrected feature set {G} will be obtained. i}
[0081] LDR image with intermediate exposure time L m The non-reference image is used as a reference image. Subsequently, histogram equalization is applied to align the non-reference image with the reference image. Relative brightness value v i The definition is as follows:
[0082]
[0083] Where h and w represent the height and width of the image, respectively. This represents the brightness value at coordinates (x, y) in the LDR image. Global brightness alignment is achieved by mapping the LDR image to the brightness range of the reference image based on the relative brightness value, as expressed by the following formula:
[0084] A i =L i ·v i i = l, m, h;
[0085] Among them, A i This indicates the corresponding brightness-aligned output. It's worth noting that, due to L... m It is a reference image, A m =L m Finally, for each frame's three inputs, we can obtain 3 pairs of 9-channel tensors {{L}. l G l A l}、{L m G m A m}、{L h G h A h}}.
[0086] Then, a set of tensors with low-exposure images as references is obtained through a swapping strategy, as expressed by the following formula:
[0087]
[0088] in This represents a tensor that is swapped from low exposure to medium exposure using a swapping strategy. This represents a tensor that is swapped from medium exposure to low exposure using a swapping strategy.
[0089] According to this exchange strategy, the present invention can use any exposure time as a reference.
[0090] Step 3: The multi-exposure image fusion HDR video generation network generates HDR video.
[0091] The HDR video generation network for multi-exposure image fusion consists of two parts: a fusion network and a restoration network. To leverage exposure time to facilitate the fusion of images with different exposure times, this invention designs an exposure time-guided attention module as the fusion network, which uses a pre-processed set of tensors referenced by the low-exposure image. {L h G h A h The image is fed into a multi-exposure image fusion HDR video generation network, which outputs the reconstructed HDR video frame by frame.
[0092] To facilitate the fusion of images with different exposure times, this invention designs an exposure time-guided attention module as a fusion network. The network structure diagram of this module is shown below. Figure 3 As shown, for a set of LDR images, the corresponding exposure time is {t}. low ,t mid ,t high The input exposure time is standardized and relativized using the following formula:
[0093]
[0094] Among them, e low and e high represents the relative exposure time for low and high exposure inputs, respectively. c is a hyperparameter, set to 10. This module employs a multi-scale structure, where the input features are downsampled three times, forming four different scales of input with the original scale input. The initial reference feature map and non-reference feature map for the j-th scale are respectively... and The input feature maps at four different scales (j = 1, 2, 3, 4, where j = 1 represents the original scale) are processed by global average pooling, two fully connected layers, and a sigmoid activation function to generate reference and non-reference feature modulation coefficients. Meanwhile, the relative exposure times of the non-reference features are each passed through a fully connected layer to obtain the time modulation coefficient t. i ∈R 1×1×128 Then, the time modulation coefficient t i The modulation coefficients are added at the channel level to the corresponding non-reference feature modulation coefficients, and then passed through three fully connected layers and one sigmoid activation function to obtain the final modulation coefficients. The final output of this module is:
[0095]
[0096] in, This serves as the output reference feature after the exposure time-guided attention module. The output non-reference features are those of the attention module guided by exposure time.
[0097] Existing attention fusion methods are used to fuse the features obtained above to obtain the output of the fusion network. Finally, an existing DomainPlus-based recovery network is used to further correct the fused features, resulting in high dynamic range (HDR) output frames. These HDR output frames are concatenated to obtain the reconstructed HDR video.
[0098] The network training steps are as follows:
[0099] L1 loss is used as the basic loss function for pixel-level supervision, and the extended advanced Sobel loss function (D-ASL) is used as an additional function, which is defined as follows:
[0100]
[0101] Where ASL is the advanced Sobel loss, and N is the training batch size. Z represents the predicted output and ground truth value of the HDR video generation network for multi-exposure image fusion, respectively, and Sobel represents the advanced Sobel filter. The expansion rate is d i The ASL is calculated using the i = {1, 2, 3} setting. Following existing techniques, this invention employs the D-ASL to formulate the loss function. The final loss function is:
[0102]
[0103] Where λ is a hyperparameter, and is set (preferably) to 0.25. Optimization is performed using the Adam optimizer, with an initial learning rate set to 10. -4 When the learning rate is below 10 -6 At this point, training ends, and the weights of each layer of the network are continuously updated by the backpropagation algorithm.
[0104] Step 4, Video Streaming:
[0105] The video stream generated by the HDR video generation network is then transmitted to the streaming module. Open-source tools such as FFmpeg are used to generate and push the video stream. The streaming module is responsible for outputting the generated HDR video in different formats, such as RAW, MPEG, and JPG, according to actual needs. This module is highly flexible, allowing the selection of suitable formats based on application requirements, and supports real-time video transmission or storage on various terminals.
[0106] On the other hand, the present invention provides an HDR video generation system for multi-exposure image fusion, including a dual-camera acquisition module, a processing unit, a video generation module, and a streaming module.
[0107] The dual-camera acquisition module is used to achieve multi-exposure image capture. The dual-camera acquisition module consists of camera 1 and camera 2, wherein:
[0108] Camera 1 (Low Exposure Camera): Used to capture high frame rate low exposure image streams to ensure that the system has a sufficient frame rate in low-light scenes to meet the requirements of video smoothness.
[0109] Camera 2 (Medium / High Exposure Camera): Alternates between capturing medium and high exposure images to complement video detail and dynamic range extension.
[0110] In low-light scenes, low-exposure image sequences are used as reference frames for generating HDR video. The dual-camera acquisition module transmits low, medium, and high-exposure image data acquired by camera 1 and camera 2 to the processing unit via a transmission interface (such as USB, WiFi, MIPI, etc.), ensuring efficient and stable data transmission.
[0111] The processing unit is used to perform frame synchronization and data preprocessing on the image data acquired by the dual-camera acquisition module.
[0112] The processing unit is first responsible for synchronizing the acquired multi-exposure image data. Since the capture of low, medium, and high-exposure images may not be synchronized, the system uses a synchronization mechanism to ensure that these images are aligned in time. Then, the synchronized multi-exposure images undergo data preprocessing, specifically including the following steps:
[0113] Gamma correction: Gamma correction is applied to low dynamic range images with different exposure times, mapping them to the high dynamic range domain to ensure uniformity of image brightness range.
[0114] Global brightness alignment: The brightness of images with different exposures is adjusted using a brightness alignment algorithm to make them consistent with the reference image.
[0115] After preprocessing, 3 pairs of 9-channel tensors (low, medium, and high exposures) are obtained. For low-light scenes, the tensors for low and medium exposures are swapped using a swapping strategy.
[0116] The video generation module uses a trained HDR video generation network based on multi-exposure image fusion to generate HDR videos.
[0117] The tensor preprocessed by the processing unit is used as input and fed into the HDR video generation network for multi-exposure image fusion to obtain reconstructed HDR frames. Multiple HDR frames are concatenated to obtain an HDR video.
[0118] The streaming module is responsible for outputting the HDR video generated by the video generation module in different formats.
[0119] In one embodiment, the dual-camera acquisition module operates as follows:
[0120] Camera 1 captures a continuous sequence of low-exposure images {L} with an exposure time of T1. l1 ,L l2 ,...,L ln Camera 2 captures medium and high exposure images alternately with exposure times T2 and T3 respectively. m ,L h}, using a low-exposure image subsequence as a reference for generating HDR video, based on the input subsequence {{L l1 ,L m ,L h},{L l2 ,L m ,L h},...,{L ln ,L m ,L h This will yield a continuous high dynamic range image subsequence output {I}. H1 ,I H2 ,...,I Hn The acquired image data consists of multiple sub-sequences. The dual-camera acquisition module transmits the obtained low, medium, and high exposure image data to the processing unit via a transmission interface.
[0121] In one embodiment, the processing unit operates as follows:
[0122] The processing unit first aligns the images transmitted from the dual-camera acquisition module in time according to their timestamps. Then, it performs data preprocessing; for each frame of video output, three LDR images with different exposure times are required. l ,L m ,L h As input, to facilitate alignment detection, the LDR image is first mapped to the HDR domain using gamma correction:
[0123]
[0124] Where γ is the gamma correction parameter, set to 2.2, t i Let G be the exposure time. Then, the gamma-corrected feature set {G} will be obtained. i}
[0125] LDR image with intermediate exposure time L m The non-reference image is used as a reference image. Subsequently, histogram equalization is applied to align the non-reference image with the reference image. Relative brightness value v i The definition is as follows:
[0126]
[0127] Where h and w represent the height and width of the image, respectively. This represents the brightness value at coordinates (x, y) in the LDR image. Global brightness alignment is achieved by mapping the LDR image to the brightness range of the reference image based on the relative brightness value, as expressed by the following formula:
[0128] A i =L i ·v i i = l, m, h;
[0129] Among them, A i This indicates the corresponding brightness-aligned output. It's worth noting that, due to L... m It is a reference image, A m =L m Finally, for each frame's three inputs, we can obtain 3 pairs of 9-channel tensors {{L}. l G l A l}、{L m G m A m}、{L h G h A h}}.
[0130] Then, a set of tensors with low-exposure images as references is obtained through a swapping strategy, as expressed by the following formula:
[0131]
[0132] in This represents a tensor that is swapped from low exposure to medium exposure using a swapping strategy.
[0133] This represents a tensor that is swapped from medium exposure to low exposure using a swapping strategy.
[0134] In one embodiment, the multi-exposure image fusion HDR video generation network consists of two parts: a fusion network and a restoration network. The fusion network is constructed based on an exposure time-guided attention module. For a set of LDR images {L... l ,L m ,L h}, and its corresponding exposure time is {t} low ,t mid ,t high After processing by the processing unit, a 9-channel tensor is obtained as the input to the fusion network. The input exposure time is standardized and relativized using the following formula:
[0135]
[0136] Among them, e low and e high represents the relative exposure time for low and high exposure inputs, respectively. c is a hyperparameter, set to 10. The exposure time-guided attention module employs a multi-scale structure, where the input features are downsampled three times sequentially, forming four different scales of input with the original scale input. The initial reference feature map and non-reference feature map at the j-th scale are respectively represented by _____. and Let j = 1, 2, 3, 4, where j = 1 represents the original scale. The four input feature maps at different scales are processed through global average pooling, two fully connected layers, and a sigmoid activation function to generate reference and non-reference feature modulation coefficients. Meanwhile, the relative exposure times of the non-reference features are each passed through a fully connected layer to obtain the time modulation coefficient t. i ∈R 1×1×128 Then, the time modulation coefficient t i Combined with the corresponding non-reference feature modulation coefficients, and passed through three fully connected layers and a sigmoid activation function, the final non-reference feature modulation coefficients are obtained. The final output of the exposure time-guided attention module is:
[0137]
[0138] in, This serves as the output reference feature after the exposure time-guided attention module. The output non-reference features are those of the attention module guided by exposure time.
[0139] An attention fusion method is used to fuse the features obtained above to obtain the output of the fusion network. Finally, an existing recovery network based on DomainPlus blocks is used to further correct the fused features to obtain high dynamic range output frames. By concatenating these high dynamic range output frames, the reconstructed HDR video can be obtained.
[0140] The network training steps are as follows:
[0141] L1 loss is used as the basic loss function for pixel-level supervision, and the extended advanced Sobel loss function (D-ASL) is used as an additional function, which is defined as follows:
[0142]
[0143] Where ASL is the advanced Sobel loss, and N is the training batch size. Z represents the predicted output and ground truth value of the HDR video generation network for multi-exposure image fusion, respectively, and Sobel represents the advanced Sobel filter. The expansion rate is d i The ASL is calculated using the i = {1, 2, 3} setting. Following existing techniques, this invention employs the D-ASL to formulate the loss function. The final loss function is:
[0144]
[0145] Where λ is a hyperparameter. Optimization is performed using the Adam optimizer, with an initial learning rate set to 10. -4 When the learning rate is below 10 -6 At this point, training ends, and the weights of each layer of the network are continuously updated by the backpropagation algorithm.
[0146] In one embodiment, the streaming module operates as follows:
[0147] The video stream generated by the HDR video generation network is then transmitted to the streaming module. Open-source tools such as FFmpeg are used to generate and push the video stream. The streaming module is responsible for outputting the generated HDR video in different formats, such as RAW, MPEG, and JPG, according to actual needs. This module is highly flexible, allowing the selection of suitable formats based on application requirements, and supports real-time video transmission or storage on various terminals.
[0148] Experimental results:
[0149] This invention trains an HDR video generation network using PyTorch on an NVIDIA RTX3090 GPU, using the training set of the Kalantari dataset, and tests it on the Kalantari test set.
[0150] Figure 5 The results of testing by simply swapping the order of low- and medium-exposure images in the Kalanitari test set show that the other two methods perform poorly after the swap, especially at high exposures. Therefore, the strategy of this invention is effective, still producing visually satisfactory results after changing the reference image, while other existing methods produce unnatural artifacts. The method of this invention can flexibly adapt to various scenarios by using images with arbitrary exposure times as a reference.
[0151] Figure 6 As can be seen from the test results in real low-light scenes, the method of the present invention can generate a high dynamic range video sequence with low-exposure images as a reference. Therefore, the frame rate of the generated video is determined by the low-exposure sequence, ensuring the real-time performance of video generation.
[0152] The above description is merely a preferred embodiment of the present invention and is not intended to limit the scope of protection of the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention are included within the scope of protection of the present invention. Parts of the present invention not described in detail are well-known to those skilled in the art.
Claims
1. A method for generating HDR video through multi-exposure image fusion, characterized in that, Includes the following steps: Step 1: Capture multi-exposure images using a dual-camera acquisition module; The dual-camera acquisition module consists of camera 1 and camera 2, wherein: Camera 1: Used to capture a high frame rate, low-exposure image stream; Camera 2: Alternately captures medium and high exposure images as a supplement to video detail and dynamic range extension; In low-light scenes, low-exposure image sequences are used as reference frames for generating HDR videos; the dual-camera acquisition module transmits low, medium, and high-exposure image data acquired by camera 1 and camera 2 to the processing unit through the transmission interface; Step 2: Frame synchronization and data preprocessing; The processing unit is first responsible for synchronizing the acquired multi-exposure image data frame by frame; then, it performs data preprocessing on the synchronized multi-exposure images, including the following preprocessing steps: Gamma correction: Perform gamma correction on low dynamic range images with different exposure times and map them to the high dynamic range domain to ensure uniformity of image brightness range; Global brightness alignment: The brightness of images with different exposures is adjusted using a brightness alignment algorithm to make them consistent with the reference image; After preprocessing, 3 pairs of 9-channel tensors with low, medium, and high exposures are obtained; for low-light scenes, the tensors with low and medium exposures are swapped using a swapping strategy. Step 3: Generate HDR video using an HDR video generation network based on multi-exposure image fusion; The preprocessed tensor is used as input and fed into the HDR video generation network for multi-exposure image fusion to obtain the reconstructed HDR frames. Multiple HDR frames are concatenated to obtain the HDR video. Step 4: Video streaming; The video stream generated by the HDR video generation network will be sent to the streaming module, which will then output the generated HDR video in different formats as needed.
2. The method for generating HDR video by multi-exposure image fusion according to claim 1, characterized in that, Step 1 is as follows: Camera 1 captures a continuous sequence of low-exposure images {L} with an exposure time of T1. l1 ,L l2 ,...,L ln Camera 2 captures medium and high exposure images alternately with exposure times T2 and T3 respectively. m ,L h }, using a low-exposure image subsequence as a reference for generating HDR video, based on the input subsequence {{L l1 ,L m ,L h },{L l2 ,L m ,L h },...,{L ln ,L m ,L h This will yield a continuous high dynamic range image subsequence output {I}. H1 ,I H2 ,...,I Hn The acquired image data consists of multiple sub-sequences. The dual-camera acquisition module transmits the low, medium, and high exposure image data to the processing unit through the transmission interface.
3. The method for generating HDR video by multi-exposure image fusion according to claim 1, characterized in that, Step 2 is as follows: The processing unit first aligns the images transmitted by the dual-camera acquisition module in time according to the timestamps; then it performs data preprocessing, requiring three LDR images with different exposure times for each frame of video output. l ,L m ,L h As input, to facilitate alignment detection, the LDR image is first mapped to the HDR domain using gamma correction: Where γ is the gamma correction parameter, set to 2.2, t i The exposure time is given; then, the gamma-corrected feature set {G} is obtained. i }; LDR image with intermediate exposure time L m As a reference image; subsequently, histogram equalization is used to align the non-reference image with the reference image; relative brightness value v i The definition is as follows: Where h and w represent the height and width of the image, respectively; This represents the brightness value at coordinates (x, y) in the LDR image. Global brightness alignment is achieved by mapping the LDR image to the brightness range of the reference image based on the relative brightness value, as expressed by the following formula: A i =L i ·v i ,i=l,m,h Among them, A i This indicates the corresponding brightness-aligned output; it is worth noting that, due to L m It is a reference image, A m =L m Finally, for each frame's three inputs, we can obtain 3 pairs of 9-channel tensors {{L}. l G l A l }、{L m G m A m }、{L h G h A h }}; Then, a set of tensors with low-exposure images as references is obtained through a swapping strategy, as expressed by the following formula: in This represents a tensor that is swapped from low exposure to medium exposure using a swapping strategy. This represents a tensor that is swapped from medium exposure to low exposure using a swapping strategy.
4. The method for generating HDR video by multi-exposure image fusion according to claim 3, characterized in that, The multi-exposure image fusion HDR video generation network consists of two parts: a fusion network and a restoration network; the fusion network is constructed based on an exposure time-guided attention module. For a set of LDR images {L l ,L m ,L h }, and its corresponding exposure time is {t} low ,t mid ,t high After processing in step 2, a 9-channel tensor is obtained as the input to the fusion network. The input exposure time is standardized and relativized using the following formula: Among them, e low and e high represents the relative exposure time of the low-exposure and high-exposure inputs, respectively; c is a hyperparameter, set to 10; the exposure time-guided attention module adopts a multi-scale structure, in which the input features are downsampled three times in sequence, forming four different scales of input with the original scale input; the initial reference feature map and non-reference feature map of the j-th scale are respectively represented by . and Let j = 1, 2, 3, 4, where j = 1 represents the original scale; the four input feature maps at different scales are processed by global average pooling, two fully connected layers, and a sigmoid activation function to generate reference and non-reference feature modulation coefficients. Meanwhile, the relative exposure times of the non-reference features are each passed through a fully connected layer to obtain the time modulation coefficient t. i ∈R 1×1×128 Then, the time modulation coefficient t i Combined with the corresponding non-reference feature modulation coefficients, and passed through three fully connected layers and a sigmoid activation function, the final non-reference feature modulation coefficients are obtained. The final output of the exposure time-guided attention module is: in, This serves as the output reference feature after the exposure time-guided attention module. The output non-reference features after the exposure time-guided attention module; An attention fusion method is used to fuse the features obtained above to obtain the output of the fusion network. Finally, an existing recovery network based on DomainPlus blocks is used to further correct the fused features to obtain high dynamic range output frames. By concatenating these high dynamic range output frames, a reconstructed HDR video can be obtained.
5. The method for generating HDR video by multi-exposure image fusion according to claim 4, characterized in that, The network training steps are as follows: L1 loss is used as the basic loss function for pixel-level supervision, and the extended advanced Sobel loss function (D-ASL) is used as an additional function, which is defined as follows: Where ASL is the advanced Sobel loss, and N is the training batch size. Z represents the predicted output and ground truth value of the HDR video generation network for multi-exposure image fusion, respectively, and Sobel represents the advanced Sobel filter. The expansion rate is d i The ASL; following existing techniques, this invention uses a setting of i = {1, 2, 3} to formulate the D-ASL; the final loss function is: Where λ is a hyperparameter; optimization is performed using the Adam optimizer, with an initial learning rate set to 10. -4 When the learning rate is below 10 -6 At this point, training ends, and the weights of each layer of the network are continuously updated by the backpropagation algorithm.
6. The method for generating HDR video by multi-exposure image fusion according to claim 5, characterized in that, λ is set to 0.
25.
7. A multi-exposure image fusion HDR video generation system, characterized in that, It includes a dual-camera acquisition module, a processing unit, a video generation module, and a streaming module; The dual-camera acquisition module is used to achieve multi-exposure image capture; the dual-camera acquisition module consists of camera 1 and camera 2, wherein: Camera 1: Used to capture a high frame rate, low-exposure image stream; Camera 2: Alternately captures medium and high exposure images as a supplement to video detail and dynamic range extension; In low-light scenes, low-exposure image sequences are used as reference frames for generating HDR videos; the dual-camera acquisition module transmits low, medium, and high-exposure image data acquired by camera 1 and camera 2 to the processing unit through the transmission interface, ensuring the efficiency and stability of data transmission. The processing unit is used to perform frame synchronization and data preprocessing on the image data acquired by the dual-camera acquisition module. The processing unit is first responsible for synchronizing the acquired multi-exposure image data frame by frame; then, it performs data preprocessing on the synchronized multi-exposure images, including the following preprocessing steps: Gamma correction: Perform gamma correction on low dynamic range images with different exposure times and map them to the high dynamic range domain to ensure uniformity of image brightness range; Global brightness alignment: The brightness of images with different exposures is adjusted using a brightness alignment algorithm to make them consistent with the reference image; After preprocessing, 3 pairs of 9-channel tensors with low, medium, and high exposures are obtained; for low-light scenes, the tensors with low and medium exposures are swapped using a swapping strategy. The video generation module uses an HDR video generation network based on multi-exposure image fusion to generate HDR videos; The tensor preprocessed by the processing unit is used as input and fed into the HDR video generation network for multi-exposure image fusion to obtain the reconstructed HDR frames. Multiple HDR frames are concatenated to obtain the HDR video. The streaming module is responsible for outputting the HDR video generated by the video generation module in different formats.