High dynamic range data generation and generalization method

By generating and migrating high dynamic range data sets, the complexity and cost of high dynamic range data acquisition in dynamic scenarios in the existing technology are solved, and large-scale and high-quality dynamic scenario HDR data generation and generalization are achieved, which improves the robustness and generalization capabilities of high dynamic range reconstruction algorithms.

CN120125484AActive Publication Date: 2025-06-10SHANGHAI ARTIFICIAL INTELLIGENCE INNOVATION CENT
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510042696.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-10
Publication Date
2025-06-10
Estimated Expiration
2045-01-10

AI Technical Summary

Technical Problem

The existing high dynamic range data acquisition methods are complex and expensive in dynamic scenarios, making it difficult to collect diversified high dynamic range dynamic data on a large scale, and the data magnitude is only about 100, which leads to visual defects such as ghosting, blurring and artifacts in large motion and high dynamic range scenarios, affecting the effect and robustness of high dynamic range reconstruction.

Method used

By collecting motion materials and high dynamic range background data, using rendering tools to process, generate rendered data sets, and adopting domain migration methods, migrating the rendered data sets to the real data sets, building a plug-and-play domain migration network structure, and achieving generalization from rendered data to real data.

Benefits of technology

It realizes batch generation of high-quality, large-scale dynamic scene HDR data, reduces labor and equipment costs, solves the distribution differences between rendered data and real data, improves the robustness and generalization capabilities of high-dynamic range reconstruction algorithms in real scenes, and reduces ghosting, fuzzing and artifact problems.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120125484A_ABST
    Figure CN120125484A_ABST
Patent Text Reader

Abstract

The invention relates to a high dynamic range data generation and generalization method, which comprises the following steps of: collecting and obtaining motion materials and high dynamic range background data, and processing the motion materials and the high dynamic range background data through a rendering tool to obtain a rendering data set, namely a high dynamic range motion data set; and generalizing the model trained on the rendering data set to a real data set by adopting a domain migration method. Compared with the prior art, on one hand, high-quality and large-scale dynamic scene high-dynamic-range data can be generated in batches, manpower and equipment costs are greatly reduced, on the other hand, the distribution difference between rendering data and real data can be solved, domain migration can be completed on real data with labels and without labels, and the data migration efficiency is improved. Therefore, the robustness and generalization ability of the high dynamic range reconstruction algorithm in a real scene are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image data processing, and in particular, to a method for generating and generalizing high-dynamic range data. Background Art

[0002] Multi-exposure synthesis is a commonly used method for shooting high-dynamic range (HDR) images, which is particularly suitable for static scenes. When shooting dynamic scenes, the movement of the moving material needs to be strictly controlled. This method relies on taking multiple photos with different exposure levels and synthesizing them into an image with a wider dynamic range. Usually, at least three images are taken, namely overexposed, normally exposed, and underexposed images. Then, specialized image processing algorithms are used to synthesize these images to capture details in different brightness ranges in the scene, and finally, a high-dynamic range image is generated. In existing research, Kalantari et al. [1] collected the first high-dynamic range data pairs for supporting high-dynamic range reconstruction algorithm research, including 74 training pairs and 15 test pairs, providing a standardized platform for training high-dynamic range reconstruction models based on deep learning methods. Next, these researchers used a similar acquisition method to collect multiple high-dynamic range data from different angles for supporting high-dynamic range reconstruction algorithm research. Tel et al. [2] collected a dataset focusing on foreground objects, environmental factors, and large motion changes, including 108 training samples and 36 test samples. Each sample captures a dynamic scene with significant foreground or camera motion. However, the motion elements and scene types are still relatively limited. Shu et al. [3] constructed a similar deghosting dataset, containing 450 training samples and 50 test samples, further improving the adaptability to dynamic scenes. With the development of mobile imaging, Kong et al. [4] captured a dataset with an extended motion range and saturation area using a smartphone, including 96 training samples and 27 test samples. However, this data acquisition method requires multiple shooting frames, so it has high requirements for the stability and performance of the camera. And when shooting a moving object scene, the movement of the moving material needs to be strictly controlled, and image alignment problems may occur between multiple exposures, resulting in artifacts or motion blur. In addition, the amount of these data is still only in the hundreds, and the data volume is still limited.

[0003] Another commonly used method is to use a beam splitter to collect high-dynamic range data. By using a beam splitter to divide the light into two beams, and then using different exposure times to shoot each beam of light separately, different exposure images are generated. These images can be combined into a high-dynamic range image. Froehlich et al. [5]Collect a batch of high-dynamic data using two beam splitters. However, beam splitters will cause some photon loss, thus affecting the image quality. And for scenes with a very large dynamic range, the beam splitter method may not be able to capture a sufficient brightness range.

[0004] In summary, the existing high-dynamic range acquisition methods for dynamic scenes rely on the control of moving materials and the setting of high-dynamic scenes, and synthesize high-dynamic range data by simultaneously capturing multiple images with different exposures. This process involves the control of the trajectory of moving objects and the synthesis of multiple exposures, with complex operations and high costs. It is difficult to collect diverse high-dynamic range dynamic data on a large scale, and the data volume is usually only about 100. Additionally, there is a method of using beam splitters for high-dynamic range data acquisition, but this method not only causes some photon loss but also cannot capture dynamic scenes with a large dynamic range. Moreover, due to the single motion type and insufficient dynamic range of the current high-dynamic range dataset, the high-dynamic range reconstruction model trained on this dataset performs poorly in dealing with the motion alignment of dynamic objects, especially in scenes with large motions and high dynamic ranges, where visual defects such as ghosting, blurring, and artifacts are likely to occur, seriously affecting the effect and robustness of high-dynamic range reconstruction. It can be said that although these above methods have promoted the development of high-dynamic range imaging, due to their limitations, especially in dynamic scenes or scenes with extremely large dynamic ranges, it is still difficult to meet the research of high-quality and high-efficiency high-dynamic range imaging algorithms.

[0005] The existing technical literature is as follows:

[0006] [1] Kalantari N K, Ramamoorthi R. Deep high dynamic range imaging of dynamic scenes[J]. ACM Transactions on Graphics(TOG), 2017, 36(4): 1 - 12.

[0007] [2] Tel S, Wu Z, Zhang Y, et al. Alignment-free HDR Deghosting with Semantics Consistent Transformer[C] / / Proceedings of the IEEE / CVF International Conference on Computer Vision. 2023: 12836 - 12845.

[0008] [3]Shu Y,Shen L,Hu X,et al.Towards Real-World HDR VideoReconstruction:A Large-Scale Benchmark Dataset and A Two-Stage AlignmentNetwork[C] / / Proceedings of the IEEE / CVF Conference on Computer Vision andPattern Recognition.2024:2879-2888.

[0009] [4]Kong L,Li B,Xiong Y,et al.SAFNet:Selective Alignment FusionNetwork for Efficient HDR Imaging[C] / / European Conference on ComputerVision.Springer,Cham,2025:256-273.

[0010] [5]Froehlich J,Grandinetti S,Eberhardt B,et al.Creating cinematicwide gamut HDR-video for the evaluation of tone mapping operators and HDR-displays[C] / / Digital photography X.SPIE,2014,9023:279-288. Summary of the Invention

[0011] The objective of the present invention is to overcome the defects of the above-mentioned existing technologies and provide a high-dynamic range data generation and generalization method, which can batch generate high-quality and large-scale dynamic scene HDR data and solve the distribution difference between rendered data and real data.

[0012] The objective of the present invention can be achieved through the following technical solutions: A high-dynamic range data generation and generalization method, comprising the following steps:

[0013] S1. Collect and obtain motion materials and high-dynamic range background data, and process them through a rendering tool to obtain a rendered dataset, that is, a high-dynamic range motion dataset;

[0014] S2. Adopt a domain transfer method to generalize the model trained on the rendered dataset to the real dataset.

[0015] Furthermore, in the step S1, the motion materials include local motion, global motion, and hybrid motion, the high-dynamic-range backgrounds include indoor backgrounds and outdoor backgrounds, and the high-dynamic-range motion datasets include indoor local motion, indoor global motion, indoor hybrid motion, outdoor local motion, outdoor global motion, and outdoor hybrid motion data.

[0016] Furthermore, the step S2 specifically realizes the generalization from rendered data to real data by constructing a plug-and-play domain transfer network structure.

[0017] Furthermore, the domain transfer network structure includes a shared branch, a transfer branch, and a pre-trained network. The shared branch is used to control the sharing of the same knowledge between the rendered data and the real data, and the transfer branch is used to control the knowledge transfer from the rendered data to the real data.

[0018] Furthermore, the shared branch consists of two layers of 1×1 convolutions. The number of input channels of the first layer of convolution is the same as the number of input channels of the pre-trained network structure, and the number of output channels is 1. The number of input channels of the second layer of convolution is 1, and the number of output channels is the same as the number of output channels of the pre-trained network structure.

[0019] Furthermore, the transfer branch consists of two layers of 1×1 convolutions. The number of input channels of the first layer of convolution is the same as the number of input channels of the pre-trained network structure, and the number of output channels is 128. The number of input channels of the second layer of convolution is 128, and the number of output channels is the same as the number of output channels of the pre-trained network structure.

[0020] Furthermore, the pre-trained network is consistent with the original network structure. The working process of the domain transfer network structure includes: inputting the input features into the pre-trained network, the shared branch, and the transfer branch respectively, then multiplying the output of the shared branch by the scaling parameter α, multiplying the output of the transfer branch by the scaling parameter β, and finally adding the scaled shared branch, the scaled transfer branch, and the output of the pre-trained network as the new output feature to replace the output feature of the original pre-trained network structure.

[0021] Furthermore, in the step S2, if there are labels in the real dataset, the domain transfer network structure is trained by the training method of gradient optimization;

[0022] If there are no labels in the real dataset, the domain transfer network structure is trained by the domain transfer framework during testing.

[0023] Furthermore, the domain transfer framework includes a data augmentation module, a student model using the domain transfer network structure, a teacher model using the domain transfer network structure, and a sliding average update weight module. The data augmentation module is used to enhance the input low-dynamic-range images.

[0024] The student model receives a low dynamic range image and obtains an output of the student model;

[0025] The teacher model receives the enhanced low dynamic range image and obtains multiple outputs of the teacher model;

[0026] The moving average weight update module updates the network weights of the teacher model by means of moving average weight update.

[0027] Further, the working process of the domain transfer framework includes: inputting the data-augmented low dynamic range image into the teacher model to obtain the output results of N teacher models, calculating the mean μ and variance σ of these N output results, where the variance σ is used to update the scaling parameters α = 1 - σ and β = 1 + σ; the mean μ is used as a pseudo label to calculate the loss function with the output result of the student network and perform gradient update to optimize the parameters of the student model;

[0028] In addition, the parameters of the teacher model are updated by means of moving average weight update, and this teacher model is used as the final inference model.

[0029] Compared with the prior art, the present invention has the following advantages:

[0030] By collecting and obtaining motion materials and high dynamic range background data, and combining with rendering tools for processing, the present invention obtains a rendered data set (i.e., a high dynamic range motion data set), and adopts a domain transfer method to transfer the rendered data set to a real data set. On the one hand, it can batch generate high-quality and large-scale dynamic scene high dynamic range data, greatly reducing the labor and equipment costs. On the other hand, it can solve the distribution difference between the rendered data and the real data, support domain transfer on labeled and unlabeled real data, thereby improving the robustness and generalization ability of the high dynamic range reconstruction algorithm in real scenes.

[0031] The present invention uses motion materials and high dynamic range background to construct a high dynamic range scene. By means of a rendering tool, a large number of motion materials are combined with the high dynamic range background, and a large-scale and high-quality high dynamic range motion data set can be efficiently rendered, including high-quality high dynamic range motion data for indoor, outdoor and various different motion types. This method can significantly improve the diversity and scale of the data set, providing richer training data for the high dynamic range reconstruction task.

[0032] The present invention designs a plug-and-play domain migration network structure, including a shared branch, a migration branch, and a pre-trained network. Among them, the shared branch is used to control the sharing of the same knowledge between the rendered data and the real data, and the migration branch is used to control the knowledge migration from the rendered data to the real data. Through the synergistic effect of these two branches, the generalization from the rendered data to the real data can be effectively achieved. In addition, when labeled real data is available, the domain migration can be directly completed through the training method of gradient optimization. When the real data has no labels, the domain migration method during testing is adopted. Through data augmentation, the student model and the teacher model based on the domain migration network structure, and combined with the method of updating weights by moving average, it is ensured that the model can still successfully migrate in the case of unlabeled data. Thus, the problem of being unable to construct a large-scale training data pair or only being able to construct a small amount of training data can be solved, and the domain migration of data can be effectively carried out regardless of whether the real data has labels or not. BRIEF DESCRIPTION OF THE DRAWINGS

[0033] Figure 1 is a schematic flowchart of the method of the present invention;

[0034] Figure 2 is a schematic diagram of constructing a high-dynamic range scene in the embodiment;

[0035] Figure 3 is a schematic diagram of high-dynamic range motion data in the embodiment;

[0036] Figure 4 is a schematic diagram of the domain migration network structure in the present invention;

[0037] Figure 5 is a schematic diagram of the working process of the domain migration framework in the present invention;

[0038] Figure 6 is a schematic diagram of the comparison of the reconstruction effects of the method of the present invention and other existing methods in the motion area;

[0039] Figure 7 is a schematic diagram of the comparison of the reconstruction effects of the method of the present invention and other existing methods in the high-dynamic range scene under direct sunlight. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0040] The present invention will be described in detail below with reference to the accompanying drawings and specific embodiments.

[0041] Embodiment

[0042] As Figure 1 shown, a high-dynamic range data generation and generalization method includes the following steps:

[0043] S1. Collect and obtain motion materials and high-dynamic range background data, and process them through a rendering tool to obtain a rendered data set, that is, a high-dynamic range motion data set;

[0044] S2. Adopt the domain adaptation method to generalize the model trained on the rendered dataset to the real dataset.

[0045] In step S1, use motion materials and high-dynamic-range backgrounds to construct a high-dynamic-range scene (as Figure 2 shown). Through the rendering tool, a large number of motion materials are combined with the high-dynamic-range background, so as to efficiently render a large-scale and high-quality high-dynamic-range motion dataset. As Figure 3 shown, the rendered dataset includes high-quality high-dynamic-range motion data for indoor, outdoor, and various different motion types. This method can significantly improve the diversity and scale of the dataset, provide richer training data for the high-dynamic-range reconstruction task, and greatly reduce the human and equipment costs by batch generating high-quality and large-scale dynamic scene high-dynamic-range data.

[0046] In step S2, in order to enable the model trained based on the rendered data to better generalize to the real data, this solution proposes a domain adaptation method that can effectively transfer the model trained on the rendered dataset to the real dataset, thereby significantly improving the performance of the model. The domain adaptation method adopts a plug-and-play design. The specific domain adaptation network structure is as Figure 4 shown, including a shared branch and a transfer branch. The shared branch is used to control knowledge sharing, while the transfer branch focuses on domain adaptation. Through the cooperation of these two branches, effective generalization from the rendered data to the real data can be achieved.

[0047] Specifically, the domain adaptation network structure includes: a shared branch, a transfer branch, and a pre-trained network structure. The shared branch is used to control the sharing of the same knowledge between the rendered data and the real data, and is composed of two layers of 1×1 convolutions. The number of input channels of the first convolution is the same as the number of input channels of the pre-trained network structure, and the number of output channels is 1. The number of input channels of the second convolution is 1, and the number of output channels is the same as the number of output channels of the pre-trained network structure;

[0048] The transfer branch is used to control the knowledge transfer from the rendered data to the real data, and is composed of two layers of 1×1 convolutions. The number of input channels of the first convolution is the same as the number of input channels of the pre-trained network structure, and the number of output channels is 128. The number of input channels of the second convolution is 128, and the number of output channels is the same as the number of output channels of the pre-trained network structure;

[0049] The pre-trained network is consistent with the original network structure.

[0050] Finally, multiply the output of the shared branch network by the scaling parameter α, and multiply the output of the migration branch network by the scaling parameter β. The initial values of the scaling parameters α and β are 1. Then, add the scaled shared branch, the scaled migration branch, and the output of the pre-trained network structure as the new output feature to replace the output feature of the original pre-trained network structure.

[0051] When labeled real data is available, domain transfer can be directly completed through a simple training method (such as gradient optimization).

[0052] When the real data has no labels, the domain transfer framework shown in Figure 5 is used to achieve domain transfer, so as to ensure that the model can still be successfully transferred and the effect can be improved in the case of unlabeled data. Specifically, the domain transfer framework includes a data augmentation module, a student model with a domain transfer structure, a teacher model with a domain transfer structure, and a moving average weight update module. Among them, the data augmentation module is used to enhance the input low-dynamic range image, including N enhancement methods such as exposure enhancement and noise enhancement. The student model with a domain transfer structure obtains the student model output after inputting the low-dynamic range image. The teacher model with a domain transfer structure obtains N teacher model outputs after inputting the enhanced low-dynamic range image. The moving average weight update module updates the network weights of the teacher model by using the method of moving average weight update.

[0053] The domain transfer process during testing is as follows: Pass the data-augmented low-dynamic range image through the teacher model to obtain the output results of N teacher models. Calculate the mean μ and variance σ of these N output results. The variance is used to update the scaling parameters α = 1 - σ and β = 1 + σ. The mean is used as the pseudo-label to calculate the loss function with the output result of the student network and perform gradient update to optimize the parameters of the student model. Then, use the method of moving average weight update to update the parameters of the teacher model, and use this teacher model as the final inference model. The whole process only needs to be executed once.

[0054] To verify the effectiveness of this solution, this embodiment compares the reconstruction effect of the method proposed in this solution with other existing methods. Among other existing methods, SCTNet is a representative method for high-dynamic range reconstruction based on the attention mechanism, and SAFNet is a representative method for high-dynamic range reconstruction using a convolutional neural network. In the SCTNet method, the dataset used is called SCT data, while in the SAFNet method, the Challenge123 dataset is used. This embodiment conducts experimental verification based on these two latest high-dynamic range reconstruction methods and their corresponding datasets. The results are as shown in Figure 6 and Figure 7 shown, as shown in Figure 6It can be seen that this solution shows the best effect in the motion area, while obvious ghosting phenomena still exist when other methods are trained on different datasets. From Figure 7 It can be seen that this solution can achieve excellent reconstruction effects in extremely high dynamic range scenarios such as direct sunlight, while other methods cannot reach the same reconstruction quality when trained in the same scenario.

[0055] Existing high dynamic range data acquisition methods are limited by the data scale, resulting in insufficient generalization ability of models trained on small-scale datasets. Using the large-scale and diverse datasets generated by this solution, the generalization and robustness of HDR reconstruction methods based on these data have been significantly improved. Especially in motion scenes and high dynamic range scenes, better reconstruction effects can be achieved.

[0056] In summary, to solve the bottleneck problem of high dynamic range reconstruction performance caused by the scarcity of dynamic scene data, this solution proposes an efficient method for generating high dynamic range training data, which can generate high-quality high dynamic range data, provide diverse dynamic scenes and rich lighting changes, and significantly increase the quantity and diversity of high dynamic range data. Using the generated rendering data, more abundant training samples can be provided in deep learning models, thus significantly enhancing the generalization ability and robustness of high dynamic range reconstruction models, effectively reducing ghosting, blurring and artifact problems in fast motion and high dynamic range scenes, and enhancing the robustness and stability of high dynamic range reconstruction models in dynamic scenes.

[0057] This solution uses the advanced graphics rendering technology of the Unreal Engine to be able to batch generate high-quality HDR data of dynamic scenes, significantly reducing the human and equipment costs. By using rendering tools, rich motion materials and dynamic range scenes are constructed, and thousands of high-quality and high dynamic range data can be generated, significantly enhancing the diversity and scale of the data. To address the distribution difference between rendering data and real data, this solution proposes a plug-and-play domain transfer method, which supports domain transfer on labeled and unlabeled real data, thereby enhancing the robustness and generalization ability of high dynamic range reconstruction algorithms in real scenes. Generalization from rendered data to real data can be achieved regardless of whether there is labeled data, and it can solve the problem of being unable to construct large-scale training data pairs or only being able to construct a small amount of training data. Effective domain transfer of data can be carried out regardless of whether the real data is labeled, especially in data-limited tasks such as high dynamic range reconstruction. This solution can significantly improve the high dynamic range imaging effects in high dynamic range and dynamic motion scenes, has high innovation and broad application prospects, and is applicable to fields such as computer vision, image enhancement, autonomous driving and virtual reality.

Claims

1. A method for generating and generalizing high dynamic range data, characterized in that: The following steps are involved: S1, collect and obtain motion materials and high dynamic range background data, and process them through a rendering tool to obtain a rendering data set, that is, a high dynamic range motion data set; S2. Use domain transfer method to generalize the model trained on the rendering dataset to the real dataset.

2. A high dynamic range data generation and generalization method according to claim 1, characterized in that: The motion material in step S1 includes local motion, global motion and mixed motion, the high dynamic range background includes indoor background and outdoor background, and the high dynamic range motion data set includes indoor local motion, indoor global motion, indoor mixed motion, outdoor local motion, outdoor global motion and outdoor mixed motion data.

3. A high dynamic range data generation and generalization method according to claim 1, characterized in that: The step S2 specifically implements generalization from rendered data to real data by constructing a plug-and-play domain migration network structure.

4. A high dynamic range data generation and generalization method according to claim 3, characterized in that: The domain migration network structure includes a sharing branch, a migration branch and a pre-trained network, the sharing branch is used to control the sharing of the same knowledge between the rendering data and the real data, and the migration branch is used to control the knowledge migration from the rendering data to the real data.

5. A high dynamic range data generation and generalization method according to claim 4, characterized in that: The shared branch consists of two layers of 1x1 convolutions. The number of input channels of the first layer of convolution is the same as that of the pre-trained network structure, and the number of output channels is 1. The number of input channels of the second layer of convolution is 1, and the number of output channels is the same as that of the pre-trained network structure.

6. A high dynamic range data generation and generalization method according to claim 4, characterized in that: The migration branch consists of two layers of 1x1 convolution. The number of input channels of the first layer of convolution is the same as that of the pre-trained network structure, and the number of output channels is 128. The number of input channels of the second layer of convolution is 128, and the number of output channels is the same as that of the pre-trained network structure.

7. A high dynamic range data generation and generalization method according to claim 4, characterized in that: The pre-trained network is consistent with the original network structure. The working process of the domain migration network structure includes: inputting input features into the pre-trained network, the shared branch and the migration branch respectively, then multiplying the output of the shared branch by the scaling parameter α, and multiplying the output of the migration branch by the scaling parameter β, and finally adding the output of the scaled shared branch, the scaled migration branch, and the pre-trained network as new output features to replace the output features of the original pre-trained network structure.

8. A high dynamic range data generation and generalization method according to claim 7, characterized in that: In step S2, if the real data set has a label, the domain transfer network structure is trained by a gradient optimization training method; If the real dataset has no labels, the domain transfer network structure is trained through the domain transfer framework during testing.

9. A high dynamic range data generation and generalization method according to claim 8, characterized in that: The domain migration framework includes a data enhancement module, a student model using a domain migration network structure, a teacher model using a domain migration network structure, and a sliding average weight update module, wherein the data enhancement module is used to enhance the input low dynamic range image; The student model receives a low dynamic range image and obtains a student model output; The teacher model receives the enhanced low dynamic range image and obtains a plurality of teacher model outputs; The sliding average weight update module adopts a sliding average weight update method to update the network weights of the teacher model.

10. A high dynamic range data generation and generalization method according to claim 9, characterized in that: The working process of the domain transfer framework includes: inputting the data-enhanced low dynamic range image into the teacher model to obtain the output results of N teacher models, calculating the mean μ and variance σ of the N output results, wherein the variance σ is used to update the scaling parameters α=1-σ, β=1+σ; the mean μ is used as a pseudo-label to calculate the loss function and perform gradient update with the output results of the student network to optimize the parameters of the student model; In addition, the parameters of the teacher model are updated by using the sliding average weight update method, and the teacher model is used as the final inference model.

Citation Information

Patent Citations

  • Image high dynamic range reconstruction method based on deep learning

    CN111292264A

  • HDR image reconstruction method based on Raw domain

    CN116757959A

  • Efficient high dynamic range visual reconstruction method and system

    CN118071663A

  • Recognition system using multimodality dataset

    US20200125902A1