Method and apparatus for constructing training data based on multi-exposure wide dynamic image
By splitting and overlaying features from wide dynamic range video RAW data, a high-quality training dataset is generated, which solves the accuracy and quality problems of wide dynamic range image processing in existing technologies and realizes intelligent and accurate image processing.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-20
- Publication Date
- 2026-03-20
AI Technical Summary
Existing technologies suffer from problems such as motion ghosting, uneven exposure, and noise interference when processing wide dynamic range images, resulting in low image processing accuracy and quality. Furthermore, traditional algorithms struggle to construct high-quality training datasets specifically for this purpose.
By acquiring target wide dynamic range video RAW data, performing data splitting and feature overlay operations, and constructing training data pairs, including pixel-level overlay and mapping of exposure linear data, a high-quality wide dynamic range image training dataset is generated.
It improves the accuracy and intelligence of image processing, enhances the image training dataset, and helps deep learning networks improve image quality in multi-exposure wide dynamic range image processing.
Smart Images

Figure CN119342353B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of data processing, and in particular to a training data construction method and device based on multi-exposure wide dynamic images. BACKGROUND
[0002] With the continuous development of science and technology, people's requirements for image quality have also increased. In the prior art, the clarity of the image is the most essential requirement of imaging technology. However, there is a difference between the dynamic range of the natural environment and the dynamic range that can be captured by the conventional linear sensor, so that in some wide dynamic scenes, the image usually has overexposure in the bright area or underexposure in the dark part.
[0003] Nowadays, most of the wide dynamic image sensors commonly used by businesses use multi-exposure methods, which usually perform image processing operations on wide dynamic images through one or more of time-sharing multiple wide dynamic, double conversion gain wide dynamic, and multi-pixel wide dynamic. However, most of the above image processing methods have processing limitations, such as image motion ghosting, poor exposure ratio linearity in some images, and incomplete spatial position matching in images. In addition, considering the noise interference under different exposure ratios, with the increasing demand for dynamic range on the product side in recent years, the traditional algorithm is increasingly limited by the design complexity and parameter tuning complexity, which hinders the effect. And the current algorithm is mainly used for single network task training, and cannot be used to deeply explore the construction of the data set, which will lead to low accuracy and low quality of image processing. Therefore, it is particularly important to provide a training data construction method for wide dynamic images to enhance the subsequent image training data set and improve the accuracy and quality of image processing. SUMMARY
[0004] The present application provides a training data construction method and device based on multi-exposure wide dynamic images, which can intelligently perform training data construction on multi-exposure wide dynamic images to obtain corresponding data pairs, which is beneficial to enhance the image training data set, and can help the deep learning network to improve the accuracy and intelligence of image processing in the actual multi-exposure wide dynamic image processing process, thereby improving the quality of image processing.
[0005] In order to solve the above technical problems, the first aspect of the present application discloses a training data construction method based on multi-exposure wide dynamic images, which comprises:
[0006] obtaining target wide dynamic video RAW data;
[0007] performing data splitting operation on the target wide dynamic video RAW data to obtain a plurality of exposure linear data;
[0008] For each of the exposure linear data, a characteristic superposition operation is performed on the exposure linear data according to a predetermined characteristic superposition parameter, to obtain exposure output data;
[0009] Based on all the exposure output data, a training data pair is constructed; wherein the training data pair includes a wide dynamic mapped image corresponding to the target wide dynamic video RAW data.
[0010] The second aspect of the present application discloses a training data construction device based on multi-exposure wide dynamic images, the device comprises:
[0011] The acquisition module is configured to acquire target wide dynamic video RAW data.
[0012] The splitting module is configured to perform a data splitting operation on the target wide dynamic video RAW data to obtain a plurality of exposure linear data.
[0013] The superposition module is configured to, for each of the exposure linear data, perform a characteristic superposition operation on the exposure linear data according to a predetermined characteristic superposition parameter, to obtain exposure output data.
[0014] The construction module is configured to construct a training data pair based on all the exposure output data; wherein the training data pair includes a wide dynamic mapped image corresponding to the target wide dynamic video RAW data.
[0015] The third aspect of the present application discloses another training data construction device based on multi-exposure wide dynamic images, the device comprises:
[0016] A memory storing executable program codes;
[0017] A processor coupled with the memory;
[0018] The processor invokes the executable program codes stored in the memory to execute the training data construction method based on multi-exposure wide dynamic images disclosed in the first aspect of the present application.
[0019] The fourth aspect of the present application discloses a computer storage medium storing computer instructions, when the computer instructions are invoked, the computer instructions are used to execute the training data construction method based on multi-exposure wide dynamic images disclosed in the first aspect of the present application.
[0020] Compared with the prior art, the embodiments of the present application have the following beneficial effects:
[0021] In this embodiment of the invention, target wide dynamic range (WDR) video RAW data is acquired, and a data splitting operation is performed on the target WDR video RAW data to obtain several exposure linear data points. For each exposure linear data point, a feature overlay operation is performed on the exposure linear data point according to a pre-determined feature overlay parameter to obtain exposure output data. Based on all exposure output data points, training data pairs are constructed. The training data pairs include the WDR-mapped image corresponding to the target WDR video RAW data. Therefore, implementing this invention can intelligently construct training data for multi-exposure WDR images to obtain corresponding data pairs, which is beneficial for enhancing the image training dataset and helping deep learning networks improve the accuracy and intelligence of image processing in actual multi-exposure WDR image processing, thereby improving the quality of image processing. Attached Figure Description
[0022] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0023] Figure 1 This is a flowchart illustrating a training data construction method based on multi-exposure wide dynamic range images disclosed in an embodiment of the present invention;
[0024] Figure 2 This is a flowchart illustrating another training data construction method based on multi-exposure wide dynamic range images disclosed in an embodiment of the present invention.
[0025] Figure 3 This is a schematic diagram of a training data construction device based on multi-exposure wide dynamic range images disclosed in an embodiment of the present invention;
[0026] Figure 4 This is a schematic diagram of another training data construction device based on multi-exposure wide dynamic range images disclosed in an embodiment of the present invention;
[0027] Figure 5 This is a schematic diagram of the structure of another training data construction device based on multi-exposure wide dynamic range images disclosed in the embodiments of the present invention;
[0028] Figure 6 This is a schematic diagram illustrating the relationship between exposure data range and actual dynamic range disclosed in an embodiment of the present invention. Detailed Implementation
[0029] In the following, the technical solutions in the embodiments of the present application will be described clearly and completely in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, rather than all the embodiments of the present application. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative effort belong to the scope of the present application.
[0030] The terms "first", "second", and the like in the description and claims of the present application and the above drawings are used to distinguish different objects, rather than to describe a particular order. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, device, product, or end including a series of steps or units is not limited to the listed steps or units, but can optionally include steps or units not listed, or can optionally include other steps or units inherent to the process, method, product, or end.
[0031] Reference herein to "an embodiment" means that a particular feature, structure, or characteristic described in connection with the embodiment can be included in at least one embodiment of the present application. The phrase appears at various places in the specification does not necessarily all refer to the same embodiment, nor is it necessarily mutually exclusive of other embodiments. It is explicitly and implicitly understood that the embodiments described herein can be combined with other embodiments.
[0032] The present application discloses a kind of based on multi-exposure wide dynamic image training data construction method and device, can be intelligently to multi-exposure wide dynamic image and execute training data construction to obtain corresponding data pair, it is favorable to enhance image training data set, and it can help depth learning network in actual multi-exposure wide dynamic image processing process improves the accuracy and intelligence of image processing, in turn, it is favorable to improve the quality of image processing. The following are described in detail respectively.
[0033] Embodiment one
[0034] Please refer to Figure 1 , Figure 1 It is a kind of based on multi-exposure wide dynamic image training data construction method flow diagram disclosed in the embodiment of the present application. Among them, Figure 1 The described multi-exposure wide dynamic image training data construction method can be applied to multi-exposure wide dynamic image training data construction device, wherein the multi-exposure wide dynamic image training data construction device can be integrated in cloud server or local server, and the embodiment of the present application is not limited. As shown in Figure 1 The multi-exposure wide dynamic image training data construction method can include the following operations:
[0035] 101、obtaining target wide dynamic video RAW data.
[0036] In the embodiment of the present application, optionally, the target wide dynamic video RAW data obtained can be collected by a high-bit-width professional dynamic camera or obtained by obtaining a public data set. The wide dynamic range (WDR) camera is an image processing technology that can simultaneously display high-light areas and dark details in the same image, solving the problem of overexposure in high-contrast scenes in traditional photography or loss of dark details. This technology realizes multi-frame exposure, image synthesis and intelligent algorithms, so that the final output image can retain details in both bright and dark parts.
[0037] In the embodiment of the present application, optionally, the public data set obtained is mainly a wide dynamic scene data set constructed by the academic circle, commonly known as HDM-HDR-2014, LiU HDRv, etc. The advantages are convenient acquisition and low cost, and the disadvantages are limited data set scenes.
[0038] In the embodiment of the present application, further optionally, the target wide dynamic video RAW data finally collected needs to be a set of continuous video data in time sequence.
[0039] 102、performing a data splitting operation on the target wide dynamic video RAW data to obtain a plurality of exposure linear data.
[0040] In the embodiment of the present application, optionally, the plurality of exposure linear data at least includes one or more of long exposure pixel values, medium exposure pixel values and short exposure pixel values of the target wide dynamic video RAW data after performing the data splitting operation.
[0041] 103、for each exposure linear data, performing a characteristic superposition operation on the exposure linear data according to a pre-determined characteristic superposition parameter to obtain exposure output data.
[0042] In the embodiment of the present application, optionally, the pre-determined characteristic superposition parameter can include one or more of a camera response curve superposition parameter, a noise superposition parameter, a motion blur superposition parameter, an inter-frame motion superposition parameter and a spatial misregistration superposition parameter.
[0043] In the embodiment of the present application, optionally, the exposure output data can include one or more of long exposure image data, medium exposure image data, short exposure image data and ultra-short exposure image data.
[0044] 104、based on all the exposure output data, constructing a training data pair.
[0045] In the embodiment of the present application, the target wide dynamic video RAW data pair in the training data pair comprises a wide dynamic mapping image corresponding to the target wide dynamic video RAW data.
[0046] In the embodiment of the present application, the training data pair can also comprise a multi-exposure image.
[0047] It can be seen that the implementation Figure 1 The described training data construction method based on multi-exposure wide dynamic image can obtain target wide dynamic video RAW data, perform data splitting operation on the target wide dynamic video RAW data to obtain a plurality of exposure linear data, perform characteristic superposition operation on the exposure linear data according to the pre-determined characteristic superposition parameter to obtain exposure output data, and construct a training data pair comprising a wide dynamic mapping image corresponding to the target wide dynamic video RAW data based on all exposure output data. The method can obtain rich image information by obtaining the target wide dynamic video RAW data, can provide the most detailed information in the image, can split the obtained RAW data into a plurality of linear data sets according to different exposure times, can obtain image data under different exposure levels, each data set highlights a specific brightness area in the image, and according to the pre-set characteristic superposition parameter, the linear data of different exposure levels are superimposed, such as pixel-level weighted superposition, to ensure that the final image can retain details in each brightness area. The exposure output data obtained through the superposition operation is paired with the original RAW data to form a training data pair, and the formed training data pair can be used for training of a machine learning model to learn how to better process wide dynamic images. In addition, by superimposing images of different exposure levels, image noise can be reduced, which is conducive to improving the overall quality of the image. Through the machine learning model, automatic wide dynamic image processing can be realized, the construction of training data for multi-exposure wide dynamic images can be intelligently performed to obtain corresponding data pairs, which is conducive to enhancing the image training data set, and the deep learning network can help improve the accuracy and intelligence of image processing in the actual multi-exposure wide dynamic image processing process, thereby improving the quality of image processing.
[0048] Embodiment two
[0049] Please refer to Figure 2 , Figure 2 is another flowchart of the training data construction method based on multi-exposure wide dynamic image disclosed in the embodiment of the present application. Among them, Figure 2 The described training data construction method based on multi-exposure wide dynamic image can be applied to a multi-exposure wide dynamic image-based training data construction device, wherein the multi-exposure wide dynamic image-based training data construction device can be integrated in a cloud server or a local server, and the embodiment of the present application is not limited. For example Figure 2As shown, the training data construction method based on the multi-exposure wide dynamic image can include the following operations:
[0050] 201, obtain target wide dynamic video RAW data.
[0051] 202, perform a pixel processing operation on the target wide dynamic video RAW data based on a predetermined pixel processing formula to obtain target linear data corresponding to the target wide dynamic video RAW data.
[0052] In the embodiment of the application, the target linear data includes rgb data after linearization processing of the target wide dynamic video RAW data.
[0053] In the embodiment of the application, optionally, the pixel processing operation can include a camera response processing operation and an electro-optical conversion function processing operation; wherein the camera response processing operation is mainly determined by the physical characteristics of the data acquisition camera, and the electro-optical conversion function processing operation is mainly determined by the photoelectric conversion characteristics defined by the physical characteristics / standards of the display device.
[0054] In the embodiment of the application, further optionally, the electro-optical conversion function can include:
[0055] ;
[0056] wherein, is the video pixel point (r / g / b data) collected / acquired, is the inverse of the electro-optical conversion function and is used to complete the conversion from the optical signal to the electrical signal, is the inverse of the camera response curve, is the rgb data after linearization processing.
[0057] 203, perform a Bayer format conversion operation on the target linear data to obtain target format data.
[0058] In the embodiment of the application, optionally, the Bayer pattern is a color filter array used for image sensors, in which each pixel point of the sensor is sensitive to only one color of light, and the usual arrangement is a 2x2 matrix, half of which is green pixels, one fourth of which is red pixels, and one fourth of which is blue pixels.
[0059] In the embodiment of the application, optionally, the above performing a Bayer format conversion operation on the target linear data to obtain target format data can include:
[0060] performing a Bayer format conversion operation on the target linear data according to a predetermined format conversion algorithm to obtain target format data;
[0061] The predetermined format conversion algorithm may include:
[0062] ;
[0063] in, This is the RGB data after prior linearization. This represents the RGB to Bayer format conversion operation, used for extracting pixels of the corresponding format at corresponding spatial locations. Bayer Data.
[0064] In this embodiment of the invention, the target format data may optionally include any one of RG / GB, BG / GR, GR / BG, and GB / RG formats.
[0065] 204. Perform a data splitting operation on the target format data to obtain several exposure linear data.
[0066] In this embodiment of the invention, optionally, the target format data can be data corresponding to any one of RG / GB, BG / GR, GR / BG, GB / RG formats; further, the target format data can be data converted to the linear domain.
[0067] 205. For each exposure linear data, perform a feature overlay operation on the exposure linear data according to the predetermined feature overlay parameters to obtain the exposure output data.
[0068] 206. Construct training data pairs based on all exposure output data.
[0069] In this embodiment of the invention, for detailed descriptions of steps 201 and 205-206, please refer to the other descriptions of steps 101 and 103-104 in Embodiment 1. These descriptions will not be repeated in this embodiment of the invention.
[0070] It is evident that implementation Figure 2The described multi-exposure wide dynamic image-based training data construction method can perform pixel processing operations on target wide dynamic video data based on a predetermined pixel processing formula to obtain corresponding target linear data, perform Bayer format conversion operations on the target linear data to obtain target format data, and then perform data splitting operations on the target format data to obtain several exposure linear data. This helps to reduce the impact of non-linear factors on image quality, lays the foundation for subsequent processing, and converts the target linear data into Bayer format data. This helps to adapt the linearized RGB data to the Bayer array of the image sensor, making subsequent image processing algorithms more effective in processing data. The data splitting operation on the converted Bayer format data obtains several exposure linear data, which can split the Bayer format data into multiple data sets according to different exposure levels, each data set highlighting a specific brightness area in the image. Through linearization processing, the non-linear distortion of the image is reduced, making the brightness and color of the image more accurate. The Bayer format conversion makes the subsequent image processing algorithm more efficient, and the data splitting operation allows the system to process data with different exposure levels separately, which helps to expand the dynamic range of the image. By processing different exposure data sets separately, the bright and dark details in the image can be better preserved, improving the overall detail performance of the image. This effectively improves the processing quality and efficiency of wide dynamic video RAW data, provides a solid foundation for generating high-quality wide dynamic images, and helps to enhance the image training data set. It can help the deep learning network improve the accuracy and intelligence of image processing in the actual multi-exposure wide dynamic image processing process, thereby improving the quality of image processing.
[0071] In an optional embodiment, performing a data splitting operation on the target format data to obtain several exposure linear data includes:
[0072] Based on the target format data, determining the exposure data range corresponding to the target wide dynamic video RAW data, and determining the exposure data relationship between the exposure data range and the actual dynamic range;
[0073] According to the exposure data relationship, determining the exposure ratio corresponding to the target format data;
[0074] According to the exposure ratio and the pre-determined splitting pixel expression, performing a data splitting operation on the target format data to obtain several exposure linear data.
[0075] In this optional embodiment, optionally, the exposure data range corresponding to the target wide dynamic video RAW data can include a long exposure acquisition range, a medium exposure acquisition range, and a short exposure acquisition range.
[0076] In the optional embodiment, optionally, the exposure ratio mainly refers to the exposure time ratio between the respective exposure images of the wide dynamic range, and more specifically, the exposure ratio here mainly refers to long exposure exposure time / middle exposure exposure time / middle exposure exposure time / short exposure exposure time.
[0077] In the optional embodiment, optionally, the exposure data relationship between the exposure data range and the actual dynamic range can include a data mapping relationship between the actually collected multi-exposure acquisition data and the actual dynamic range.
[0078] In the optional embodiment, further optionally, the exposure data relationship between the exposure data range and the actual dynamic range can be as shown in Figure 6 Figure 6 is a schematic diagram of an exposure data relationship between an exposure data range and an actual dynamic range disclosed by the embodiment of the application, as shown in Figure 6 , taking 3-frame wide dynamic image acquisition as an example, denoting W(L), W(M), and W(S) as the actual scene dynamics recorded by long (Long Time Exposure) / middle (Middle Time Exposure) / short (Short Time Exposure) exposure, respectively, then the dynamic range W(H) of the 3-frame wide dynamic actual recording is the union of the three, which can be expressed as W(H)=W(L)∪W(M)∪W(S), and ∪ represents the mathematical union. Further, corresponding to the actual response pixel, taking 3-frame 16 times exposure ratio wide dynamic (single exposure 12-bit linear ADC conversion) as an example, the corresponding split pixel value can be expressed as:
[0079]
[0080] Among them, is an ideal wide dynamic pixel value (in the present application, the data after the previous mosaic processing ), , , corresponding to the long exposure pixel value, the middle exposure pixel value, and the short exposure pixel value after splitting.
[0081] In the optional embodiment, optionally, the exposure ratio mainly refers to the exposure time ratio between the respective exposure images of the wide dynamic range, and more specifically, the exposure ratio here mainly refers to long exposure exposure time / middle exposure exposure time / middle exposure exposure time / short exposure exposure time; single exposure 12-bit linear ADC conversion mainly refers to the conversion accuracy of analog signal conversion to digital signal in the image sensor data acquisition process. The bit width of the industry scene CMOS is 10 / 12 / 14 bits, and the wider the bit width, the wider the dynamic range recorded, and the higher the cost / power consumption.
[0082] It can be seen that implementing the optional embodiment can determine the exposure data range corresponding to the target wide dynamic video RAW data based on the target format data, determine the exposure data relationship between the exposure data range and the actual dynamic range, determine the exposure ratio corresponding to the target format data according to the exposure data relationship, and perform data splitting operation on the target format data according to the exposure ratio and the split pixel expression to obtain a plurality of exposure linear data. Based on the target format data, the exposure data range of the original wide dynamic video RAW data can be determined, the data of the brightest and darkest parts in the image can be identified, the corresponding relationship between the exposure data range and the actually observable dynamic range can be analyzed and determined, which helps to understand how the data under different exposure levels is mapped to the brightness change that can be perceived by the human eye. According to the exposure data relationship, the exposure ratio of the target format data is calculated, that is, how the data under different exposure levels is proportionally mapped to the final image. By using the exposure ratio and the pre-determined split pixel expression, the target format data is split to obtain a plurality of linear data sets under different exposure levels. The dynamic range of the image can be expanded by precisely controlling the data under different exposure levels. The dynamic range of the image can be expanded by precisely controlling the data under different exposure levels, so that the image details of the highlight and low-illumination regions can be preserved, and the data splitting under different exposure levels can reveal more details in the image, which is beneficial to improve the data amount of the image data set. By pre-calculating the exposure ratio and the split pixel expression, the complexity of post-processing can be reduced, the efficiency of image processing can be improved, and image noise and artifacts can be reduced. Especially when synthesizing data under different exposure levels, the transition between different regions can be smoother. By dynamically adjusting the exposure ratio, different lighting environments can be adapted to ensure high-quality images under various conditions. Through precise exposure control and data splitting, the processing effect of wide dynamic video RAW data is significantly improved, which is beneficial to enhance the image training data set and help the deep learning network improve the accuracy and intelligence of image processing in the actual multi-exposure wide dynamic image processing process, thereby improving the quality of image processing.
[0083] In another optional embodiment, for each exposure linear data, a characteristic superposition operation is performed on the exposure linear data according to the pre-determined characteristic superposition parameter to obtain exposure output data, including:
[0084] For each exposure linear data, a first characteristic superposition operation is performed on the exposure linear data to obtain first linear data corresponding to the exposure linear data.
[0085] For each exposure linear data, a second characteristic superposition operation is performed on the first linear data corresponding to the exposure linear data to obtain second linear data corresponding to the exposure linear data.
[0086] For each exposure linear data, exposure output data is generated according to second linear data corresponding to the exposure linear data;
[0087] The first characteristic superposition operation includes a camera response curve superposition operation and a noise superposition operation, and the second characteristic superposition operation includes a motion blur superposition operation, an interframe motion superposition operation, and a spatial misregistration superposition operation.
[0088] In the optional embodiment, the camera response curve superposition operation mainly considers the difference of the image sensor response curves at different ISOs, and the multi-exposure wide dynamic image sensor is actually at different ISOs in the actual image acquisition process.
[0089] It can be seen that by implementing the optional embodiment, the first characteristic superposition operation can be performed on each exposure linear data to obtain the first linear data corresponding to the exposure linear data, and the second characteristic superposition operation can be performed on the first linear data corresponding to each exposure linear data to obtain the corresponding second linear data, and the exposure output data can be generated according to the second linear data, wherein the first characteristic superposition operation includes camera response curve superposition operation and noise superposition operation; the second characteristic superposition operation includes motion blur superposition operation, inter-frame motion superposition operation and spatial misregistration superposition operation, and the exposure output data includes long-exposure pixel value data, medium-exposure pixel value data and short-exposure pixel value data corresponding to the exposure linear data. The pixel value can be adjusted by the first characteristic superposition operation to match the actual response characteristics of the camera and the noise superposition can be performed on the exposure linear data to simulate and reduce image noise. The motion blur superposition can be performed on the first linear data corresponding to each exposure linear data by the second characteristic superposition operation to simulate the blurring effect of moving objects on the exposure device, the inter-frame motion superposition can be performed to handle the motion changes between consecutive frames to improve the image stability in dynamic scenes, and the spatial misregistration superposition can be performed to adjust the pixel position to simulate the actual motion and scene changes. By simulating the camera response and noise characteristics, the generated image is closer to the actual shooting effect. The motion blur and inter-frame motion superposition operations can help better handle dynamic scenes and reduce motion artifacts. The spatial misregistration superposition operation can improve the stability and continuity of the image in fast motion or complex scenes. By processing the long-exposure, medium-exposure and short-exposure pixel value data respectively, better exposure balance can be achieved in different lighting conditions, which can adapt to different shooting environments and lighting conditions to provide consistent image quality. Through detailed characteristic superposition operations, not only the quality and authenticity of the image are improved, but also the processing capability of dynamic scenes is enhanced, so that the final exposure output data can maintain high quality under various conditions. Through precise exposure control and data splitting, the processing effect of wide dynamic video RAW data is significantly improved, which is conducive to enhancing the image training data set and helping the deep learning network to improve the accuracy and intelligence of image processing in the actual multi-exposure wide dynamic image processing process, thereby improving the quality of image processing.
[0090] In yet another optional embodiment, based on all exposure output data, a training data pair is constructed, including:
[0091] performing a first ground truth data construction operation on the exposure output data based on the exposure output data and a predetermined camera response curve to obtain first ground truth construction data, wherein the first ground truth construction data comprises a long-exposure pixel value after superimposing the exposure output data on the camera response curve, a middle-exposure pixel value after superimposing the exposure output data on the camera response curve, a short-exposure pixel value after superimposing the exposure output data on the camera response curve, and an ultra-short-exposure pixel value after superimposing the exposure output data on the camera response curve;
[0092] performing a second ground truth data construction operation on the first ground truth construction data based on the first ground truth construction data and a predetermined noise processing parameter to obtain second ground truth construction data, wherein the second ground truth construction data comprises a middle-exposure pixel value after superimposing noise on the first ground truth construction data, a short-exposure pixel value after superimposing noise on the first ground truth construction data, and an ultra-short-exposure pixel value after superimposing noise on the first ground truth construction data;
[0093] generating target training data based on the first ground truth construction data and the second ground truth construction data, in combination with a predetermined inter-frame motion superimposition ground truth data construction formula and a predetermined spatial misregistration superimposition ground truth data construction formula;
[0094] obtaining a training data pair according to all the target training data.
[0095] In this optional embodiment, optionally, the above-mentioned performing a first ground truth data construction operation on the exposure output data based on the exposure output data and a predetermined camera response curve to obtain first ground truth construction data can comprise:
[0096] performing a first ground truth data construction operation on the exposure output data based on the exposure output data and a predetermined camera response curve to obtain first ground truth construction data;
[0097] wherein the predetermined camera response curve corresponds to a ground truth data construction algorithm, and the ground truth data construction algorithm comprises:
[0098] ;
[0099] wherein, is a long-exposure BAYER pixel value / middle-exposure BAYER pixel value / short-exposure BAYER pixel value / ultra-short-exposure BAYER pixel value obtained in the preceding dynamic image splitting, represents a camera response curve, is a long-exposure pixel value / middle-exposure pixel value / short-exposure pixel value / ultra-short-exposure pixel value after superimposing the camera response curve;
[0100] Further, considering that low ISO usually has better linearity, it is suggested to use the camera response curve under low ISO; and, the biggest difference between this and the previous data construction is that The curve does not change with the camera ISO, which is used as the true value construction process.
[0101] In this optional embodiment, optionally, the above predetermined noise processing parameters can include:
[0102] ;
[0103] wherein, is the pixel value of the middle exposure / short exposure / ultra-short exposure after superimposing the camera response curve of the prequel, represents the noise curve of long exposure, is the pixel value of the middle exposure / short exposure / ultra-short exposure after superimposing the noise; further, It can be obtained by light box calibration / noise model modeling, and the true value data corresponding to the noise superposition can maintain the same signal-to-noise ratio as the long exposure.
[0104] In this optional embodiment, optionally, the predetermined inter-frame motion superposition true value data construction formula can include:
[0105] ;
[0106] wherein, , , represents the long exposure image / middle exposure image / short exposure image / ultra-short exposure image after reconstructing the time sequence; further, as the true value data for network training, there is no misalignment in the time sequence.
[0107] In this optional embodiment, optionally, the predetermined spatial misalignment superposition true value data construction formula can include:
[0108] ;
[0109] wherein, , , , represents the long exposure image / middle exposure image / short exposure image / ultra-short exposure image after spatial sampling reconstruction; further, as the true value data for network training, there is no misalignment in the spatial position.
[0110] In this optional embodiment, optionally, the above training data pair obtained according to all target training data can include:
[0111] All target training data is determined as training data pairs.
[0112] It can be seen that implementing this optional embodiment can perform a first ground truth data construction operation based on exposure output data and camera response curve to obtain first ground truth construction data, and perform a second ground truth data construction operation based on first ground truth construction data and noise processing parameters to obtain second ground truth construction data, generate target training data based on first ground truth data construction and second ground truth data construction combined with inter-frame motion superimposed ground truth data construction formula and spatial misregistration superimposed ground truth data construction formula, and then obtain training data pairs, which can combine long exposure, medium exposure and short exposure pixel value data with camera response curve through exposure output data and camera response curve to simulate the real response when the camera captures an image, and obtain second ground truth construction data based on first ground truth construction data and noise processing parameters, which can increase the complexity of the real world to the training data by simulating the noise in actual shooting, generate target training data by combining first ground truth construction data, second ground truth construction data, inter-frame motion superimposed ground truth data construction formula and spatial misregistration superimposed ground truth data construction formula, which helps to simulate motion blur and spatial changes in dynamic scenes and provides more comprehensive training information for the algorithm, and the algorithm can more accurately process dynamic scenes through data construction containing inter-frame motion and spatial misregistration, the algorithm can learn how to more effectively identify and reduce image noise by superimposing noise in training data, the adaptability of the model to different shooting conditions can be improved through training data pairs containing various complex situations, the authenticity and diversity of training data pairs help to improve the accuracy of image analysis tasks such as object detection and image segmentation, and the performance and robustness of image processing algorithms can be significantly improved through the construction of training data pairs containing various factors and complexity, especially when processing real-world image data, and then the processing effect of wide dynamic video RAW data can be significantly improved through precise exposure control and data splitting, which is beneficial to enhance the image training data set, and can help the deep learning network to improve the accuracy and intelligence of image processing in the actual multi-exposure wide dynamic image processing process, and then it is beneficial to improve the quality of image processing.
[0113] In yet another optional embodiment, based on all exposure output data, training data pairs are constructed, including:
[0114] Based on the exposure output data and the pre-determined inverse camera response curve, wide dynamic linear ground truth data is determined;
[0115] Based on the wide dynamic linear ground truth data, training data pairs are constructed;
[0116] And, after determining the wide dynamic linear ground truth data based on the exposure output data and the pre-determined inverse camera response curve, the method further includes:
[0117] performing a tone mapping operation on the wide dynamic linear ground truth data according to a predetermined tone mapping algorithm to obtain tone mapping ground truth data;
[0118] The training data pair is constructed based on the wide dynamic linear ground truth data, and includes:
[0119] The training data pair is constructed based on the tone mapping ground truth data.
[0120] In this optional embodiment, optionally, the ground truth of the wide dynamic linear data generated by the exposure output data and the predetermined inverse camera response curve can use the wide dynamic linear data after the inverse camera response curve of the wide dynamic video collected as the ground truth, and if there is inter-frame data sampling / space domain data sampling due to the characteristics of the image sensor, the wide dynamic image sampling can be performed according to the two parts of the foregoing multiple descriptions of inter-frame motion superposition / space misplacement superposition, wherein the following algorithm can be used for description:
[0121] ;
[0122] ;
[0123] wherein, corresponding to inter-frame motion superposition, corresponding to space misplacement superposition, , , , respectively correspond to the wide dynamic images before and after sampling, and the subscript , and , respectively correspond to the frame number in inter-frame sampling and the horizontal and vertical coordinate positions in space sampling.
[0124] In this optional embodiment, optionally, the predetermined tone mapping algorithm can include:
[0125] ;
[0126] wherein, is the ground truth data obtained in the data construction, is a preset tone mapping operator, represents the final tone mapping ground truth data set; wherein the preset tone mapping operator can be a traditional tone mapping algorithm, or a tone mapping deep learning network.
[0127] It can be seen that implementing the optional embodiment can determine wide dynamic linear true value data based on exposure output data and inverse camera response curve, construct training data pairs based on wide dynamic linear true value data, and also perform tone mapping operation on wide dynamic linear true value data according to tone mapping algorithm after determining wide dynamic linear true value data to obtain tone mapping true value data, construct training data pairs based on tone mapping true value data, determine wide dynamic linear true value data using exposure output data and inverse camera response curve, convert camera-captured data to linear space for better processing of wide dynamic range images, and use the determined wide dynamic linear true value data to construct data pairs for training image processing models, and further construct training data pairs by using tone-mapped data to train the model's color and brightness performance under different display conditions. Through the processing of the inverse camera response curve, the details of the high dynamic range scene can be better restored. The use of the tone mapping algorithm helps to ensure the consistency and accuracy of colors on different display devices. The constructed training data pairs contain the complexity of wide dynamic range and color mapping, which helps to train a more robust and accurate model. By incorporating tone mapping true value data into the training process, the tone mapping algorithm can be further optimized to adapt to different visual and display requirements. Through the construction of training data pairs, the model can learn how to present the best image effect on different display devices, thereby providing rich training information for image processing models by combining wide dynamic linear true value data and tone mapping true value data, thereby achieving significant performance improvement in wide dynamic range imaging and color management, and thus significantly improving the processing effect of wide dynamic video RAW data, which is conducive to enhancing the image training data set and helping the deep learning network to improve the accuracy and intelligence of image processing in the actual multi-exposure wide dynamic image processing process, thereby improving the quality of image processing.
[0128] In yet another optional embodiment, for each exposure linear data, a first characteristic superposition operation is performed on the exposure linear data to obtain first linear data corresponding to the exposure linear data, including:
[0129] For each exposure linear data, a first characteristic superposition operation is performed on the exposure linear data according to a pre-determined camera response curve superposition algorithm and a pre-determined noise superposition algorithm to obtain first linear data corresponding to the exposure linear data;
[0130] The pre-determined camera response curve superposition algorithm includes:
[0131] ;
[0132] wherein, the long-exposure Bayer pixel value / medium-exposure Bayer pixel value / short-exposure Bayer pixel value / super-short-exposure Bayer pixel value obtained in the preceding dynamic image splitting, (.) the camera response curve under each illumination, the long-exposure pixel value / medium-exposure pixel value / short-exposure pixel value / super-short-exposure pixel value after superimposing the camera response curve;
[0133] wherein the pre-determined noise superimposition algorithm comprises:
[0134] ;
[0135] wherein, the long-exposure pixel value / medium-exposure pixel value / short-exposure pixel value / super-short-exposure pixel value after superimposing the camera response curve in the preceding, the noise curve under each illumination, the long-exposure pixel value / medium-exposure pixel value / short-exposure pixel value / super-short-exposure pixel value after superimposing the noise;
[0136] and for each exposure linear data, performing a second characteristic superimposition operation on the first linear data corresponding to the exposure linear data to obtain the second linear data corresponding to the exposure linear data, comprising:
[0137] for each exposure linear data, performing a second characteristic superimposition operation on the first linear data corresponding to the exposure linear data according to the pre-determined motion blur superimposition algorithm, inter-frame motion superimposition algorithm and spatial misregistration superimposition algorithm to obtain the second linear data corresponding to the exposure linear data;
[0138] wherein the motion blur superimposition algorithm comprises:
[0139] ;
[0140] wherein, the long-exposure image / medium-exposure image after superimposing the motion blur, the long-exposure image / medium-exposure image Bayer pixel input image after superimposing the noise in the preceding, a spatial domain blur kernel function that varies with exposure;
[0141] the inter-frame motion superimposition algorithm comprises:
[0142] ;
[0143] wherein, denotes the long-exposure image / medium-exposure image / short-exposure image / super-short-exposure image in time sequence, the subscript thereof representing the corresponding frame number on the time axis; , , representing the long exposure image / medium exposure image / short exposure image / ultra-short exposure image after reconstruction of the time sequence;
[0144] spatial dislocation superposition algorithm, comprising:
[0145] ;
[0146] wherein, , , , representing the long exposure image / medium exposure image / short exposure image / ultra-short exposure image before sampling, wherein the subscript represents the spatial position coordinates of the corresponding pixel points, , , , representing the long exposure image / medium exposure image / short exposure image / ultra-short exposure image after spatial domain sampling reconstruction.
[0147] In this optional embodiment, the camera response curve superposition is used to consider the difference of the image sensor response curve at different ISOs, and in the actual image acquisition process of the multi-exposure wide dynamic image sensor, each exposure is actually at different ISOs.
[0148] In this optional embodiment, the noise superposition algorithm is used for the multi-exposure wide dynamic image sensor in the actual image acquisition process, and each exposure is actually at different ISOs, and the image noise level at different ISOs is greatly different according to the actual image sensor characteristics.
[0149] In this optional embodiment, further optionally, the motion blur superposition operation is mainly aimed at the case that the actual physical exposure time corresponding to each exposure in the multi-exposure wide dynamic image sensor is not consistent, and the clarity of the moving object corresponding to each exposure is not consistent; for example, the short exposure image can be taken as the reference (without blur), and the moving object in the corresponding medium exposure image / long exposure image can be blurred by different degrees of motion blur kernel. Further, the motion blur superposition operation can be realized through a motion blur superposition algorithm;
[0150] wherein, the motion blur superposition algorithm can comprise:
[0151] ;
[0152] wherein, is the long exposure image / medium exposure image after motion blur superposition, is the long exposure image / medium exposure image Bayer pixel input image after the preceding noise superposition, is a spatial blur kernel function that varies with exposure, where the longer the exposure time, the stronger the effect of the corresponding post-function, further, the function can be adjusted by adjusting the blur function parameters or blur radius, and the process can be described as a Gaussian blur kernel function, where the Gaussian blur kernel function can include:
[0153] ;
[0154] wherein, represents the starting and ending coordinates of the filter in the vertical direction, , respectively represent the weight and pixel at the corresponding spatial position of the filter, represents the spatial position coordinates of the pixel; further, The calculation process of can be calculated by the following function:
[0155] ;
[0156] wherein, represents the standard deviation of the Gaussian distribution, wherein, affects the final filtering strength.
[0157] In this optional embodiment, further optionally, the inter-frame motion superposition operation is specifically used for the case where multiple exposures come from the same physical pixel, and there is a misalignment of moving objects between exposures during the acquisition of moving objects; further, the misalignment of inter-frame motion of the moving objects can be simulated by inter-frame motion superposition, which is realized by sampling combination between sequence frames without motion misalignment, for example, taking a 3-frame wide dynamic as an example, the implementation process can be described as follows:
[0158] ;
[0159] wherein, , , represents a long exposure image / a medium exposure image / a short exposure image / an ultra-short exposure image in a time sequence, and the subscript represents the corresponding frame number on the time axis; , , represents a reconstructed long exposure image / a medium exposure image / a short exposure image / an ultra-short exposure image in a time sequence. As can be seen from the above formula, the number of frames in the newly constructed long exposure time sequence / medium exposure time sequence / short exposure time sequence / ultra-short exposure time sequence is only 1 / 3 of the original data set, and the exposure sequence is long exposure / medium exposure / short exposure / ultra-short exposure (preferably long exposure first, then medium exposure, and finally short exposure).
[0160] In the optional embodiment, further optionally, the spatial misalignment superposition operation is mainly for the case that the physical pixels corresponding to each exposure correspond to misalignment in space; wherein in the actual wide dynamic image acquisition process, the four are set to different exposure times and gains to simultaneously perform image acquisition, the pixel points between each exposure have a certain degree of spatial position misalignment, and such misalignment will produce problems such as jaggies at the texture edges of the actual image. The superposition of spatial misalignment mainly expects to simulate this case and enable the network to have the error correction capability of such data, and the main means is to combine through spatial position sampling of the original sequence, and the implementation process can be described as follows:
[0161]
[0162] wherein, , , , represents the long exposure image / medium exposure image / short exposure image / super short exposure image before sampling, and the subscript represents the spatial position coordinates of the corresponding pixel points, , , , represents the long / medium / short / super short exposure image after spatial sampling reconstruction, and from the above formula, it can be seen that the spatial resolution of the newly constructed long exposure image / medium exposure image / short exposure image / super short exposure image is only 1 / 4 of that before sampling, and in turn, the long exposure image / medium exposure image / short exposure image / super short exposure image.
[0163] It can be seen that implementing the optional embodiment can combine the predetermined camera response curve superposition algorithm and the noise superposition algorithm to perform the first characteristic superposition operation on the first linear data corresponding to the exposure linear data, can generate an image closer to the real shooting effect by simulating the actual response and noise characteristics of the camera, and the superposition of the camera response curve helps to expand the dynamic range of the image, better preserves the details of the highlight and shadow parts, and the superposition of the noise model helps to consider the noise influence in the actual shooting in the image processing algorithm, thereby improving the effect of noise reduction processing, the training data pair contains camera response and noise information, which helps to improve the accuracy and reliability of the image processing algorithm in actual application, the characteristic superposition operation can ensure the consistency and naturalness of the synthesis result when synthesizing multi-exposure high dynamic range images, and by considering the camera response and noise characteristics under different ISO settings, the method can adapt to different shooting environments and lighting conditions, more real and accurate training data pairs in image analysis and computer vision tasks can improve the accuracy of task execution, automatic image processing flow can utilize these characteristic superposition operations to reduce manual intervention, improve processing efficiency and consistency, thereby the characteristic superposition operation can simulate the response and noise characteristics of the camera to provide more rich and real training data for image processing, thereby improving the performance and quality of image processing in multiple aspects, thereby significantly improving the processing effect of wide dynamic video RAW data, which is beneficial to enhancing the image training data set, and can help the deep learning network to improve the accuracy and intelligence of image processing in the actual multi-exposure wide dynamic image processing process, thereby being beneficial to improving the quality of image processing.
[0164] Embodiment three
[0165] Please refer to Figure 3 , Figure 3 is a structure schematic view of a training data construction device based on multi-exposure wide dynamic image disclosed by the embodiment of the application. As Figure 3 shown, the training data construction device based on multi-exposure wide dynamic image can include:
[0166] The acquisition module 301 is configured to acquire target wide dynamic video RAW data.
[0167] The splitting module 302 is configured to perform data splitting operation on the target wide dynamic video RAW data to obtain a plurality of exposure linear data.
[0168] The superposition module 303 is configured to, for each exposure linear data, perform characteristic superposition operation on the exposure linear data according to the predetermined characteristic superposition parameters to obtain exposure output data.
[0169] The construction module 304 is configured to construct a training data pair based on all the exposure output data, wherein the training data pair includes a wide dynamic mapping image corresponding to the target wide dynamic video RAW data.
[0170] It can be seen that the implementation Figure 3 The described device can obtain target wide dynamic video RAW data, perform data splitting operation on the target wide dynamic video RAW data to obtain a plurality of exposure linear data, perform characteristic superposition operation on the exposure linear data according to the pre-determined characteristic superposition parameter to obtain exposure output data, and construct a training data pair including a wide dynamic mapping image corresponding to the target wide dynamic video RAW data based on all the exposure output data. The device can obtain rich image information by obtaining the target wide dynamic video RAW data, can provide the most detailed information in the image, can split the obtained RAW data into a plurality of linear data sets according to different exposure times, can obtain image data under different exposure levels, each data set highlights a specific brightness area in the image, and according to the pre-set characteristic superposition parameter, the linear data of different exposure levels are superimposed, such as pixel-level weighted superposition, to ensure that the final image can retain details in each brightness area. The exposure output data obtained through the superposition operation is paired with the original RAW data to form a training data pair, and the formed training data pair can be used for training of a machine learning model to learn how to better process wide dynamic images. In addition, by superimposing images of different exposure levels, image noise can be reduced, which is conducive to improving the overall quality of the image. Through the machine learning model, automatic wide dynamic image processing can be realized, and the construction of the training data of the multi-exposure wide dynamic image can be intelligently performed to obtain the corresponding data pair, which is conducive to enhancing the image training data set. In addition, the deep learning network can help improve the accuracy and intelligence of image processing in the actual multi-exposure wide dynamic image processing process, thereby improving the quality of image processing.
[0171] In an optional embodiment, as Figure 4 shown, the device further includes:
[0172] The processing module 305 is configured to perform pixel processing operation on the target wide dynamic video RAW data based on a pre-determined pixel processing formula to obtain target linear data corresponding to the target wide dynamic video RAW data after the acquisition module 301 acquires the target wide dynamic video RAW data and before the splitting module 302 performs data splitting operation on the target wide dynamic video RAW data to obtain a plurality of exposure linear data, wherein the target linear data includes rgb data after linearization processing of the target wide dynamic video RAW data.
[0173] The conversion module 306 is configured to perform Bayer format conversion operation on the target linear data to obtain target format data.
[0174] The specific manner in which the splitting module 302 performs the data splitting operation on the target wide dynamic video RAW data to obtain the plurality of exposure linear data includes:
[0175] performing the data splitting operation on the target format data to obtain the plurality of exposure linear data.
[0176] It can be seen that the implementation Figure 4 The described apparatus can perform a pixel processing operation on target wide dynamic video data based on a predetermined pixel processing formula to obtain corresponding target linear data, perform a Bayer format conversion operation on the target linear data to obtain target format data, and then perform a data splitting operation on the target format data to obtain a plurality of exposure linear data. This can help to reduce the impact of non-linear factors on image quality, lay a foundation for subsequent processing, and convert the target linear data into Bayer format data. This can help the linearized RGB data to be adapted to the Bayer array of the image sensor, making the subsequent image processing algorithm more effective in processing data. The data splitting operation on the converted Bayer format data can split the Bayer format data into a plurality of data sets according to different exposure levels, with each data set highlighting a specific brightness area in the image. Through linearization processing, the non-linear distortion of the image is reduced, making the brightness and color of the image more accurate. The Bayer format conversion makes the subsequent image processing algorithm more efficient, and the data splitting operation allows the system to process data of different exposure levels separately, which helps to expand the dynamic range of the image. By processing different exposure data sets separately, the bright and dark details in the image can be better preserved, improving the overall detail performance of the image. This can effectively improve the processing quality and efficiency of wide dynamic video RAW data, providing a solid foundation for generating high-quality wide dynamic images, which is conducive to enhancing image training data sets and helping deep learning networks improve image processing accuracy and intelligence in the actual multi-exposure wide dynamic image processing process, thereby improving image processing quality.
[0177] In another optional embodiment, as Figure 4 shown, the splitting module 302 performs a data splitting operation on the target format data to obtain a plurality of exposure linear data, and the specific process includes:
[0178] determining, based on the target format data, an exposure data range corresponding to the target wide dynamic video RAW data, and determining an exposure data relationship between the exposure data range and an actual dynamic range;
[0179] determining, according to the exposure data relationship, an exposure ratio corresponding to the target format data;
[0180] According to the exposure ratio and the predetermined split pixel expression, a data split operation is performed on the target format data to obtain a plurality of exposure linear data.
[0181] It can be seen that the implementation Figure 4 The described device can determine the exposure data range corresponding to the target wide dynamic video RAW data based on the target format data, determine the exposure data relationship between the exposure data range and the actual dynamic range, determine the exposure ratio corresponding to the target format data according to the exposure data relationship, and perform a data split operation on the target format data according to the exposure ratio and the split pixel expression to obtain a plurality of exposure linear data. Based on the target format data, the exposure data range of the original wide dynamic video RAW data can be determined, the data of the brightest and darkest parts of the image can be identified, the corresponding relationship between the exposure data range and the actually observable dynamic range can be analyzed and determined, which helps to understand how the data at different exposure levels is mapped to the brightness change that can be perceived by the human eye. According to the exposure data relationship, the exposure ratio of the target format data is calculated, i.e. how the data at different exposure levels is proportionally mapped to the final image. By using the exposure ratio and the predetermined split pixel expression, the target format data is split to obtain a plurality of linear data sets at different exposure levels. The dynamic range of the image can be expanded by precisely controlling the data at different exposure levels. The dynamic range of the image can be expanded by precisely controlling the data at different exposure levels, so that the image details of high light and low illumination areas can be preserved, and the data split at different exposure levels can reveal more details in the image, which is beneficial to improve the data amount of the image data set. By pre-calculating the exposure ratio and the split pixel expression, the complexity of post-processing can be reduced, the efficiency of image processing can be improved, and image noise and artifacts can be reduced. In particular, when different exposure level data is synthesized, different regions can be smoothly transitioned. By dynamically adjusting the exposure ratio, different lighting environments can be adapted to ensure high-quality images under various conditions. Through precise exposure control and data split, the processing effect of wide dynamic video RAW data is significantly improved, which is beneficial to enhance the image training data set and help the deep learning network improve the accuracy and intelligence of image processing in the actual multi-exposure wide dynamic image processing process, thereby improving the quality of image processing.
[0182] In yet another optional embodiment, as Figure 4 shown, the superposition module 303 performs a characteristic superposition operation on each exposure linear data according to the predetermined characteristic superposition parameter to obtain the specific manner of the exposure output data, which includes:
[0183] For each exposure linear data, a first characteristic superposition operation is performed on the exposure linear data to obtain the first linear data corresponding to the exposure linear data.
[0184] for each exposure linear data, performing a second characteristic superposition operation on the first linear data corresponding to the exposure linear data to obtain the second linear data corresponding to the exposure linear data;
[0185] for each exposure linear data, generating exposure output data according to the second linear data corresponding to the exposure linear data;
[0186] wherein the first characteristic superposition operation includes camera response curve superposition operation, noise superposition operation; the second characteristic superposition operation includes motion blur superposition operation, inter-frame motion superposition operation, spatial misregistration superposition operation; the exposure output data corresponding to each exposure linear data includes long-exposure pixel value data, medium-exposure pixel value data, short-exposure pixel value data, and ultra-short-exposure pixel value data corresponding to the exposure linear data.
[0187] It can be seen that the implementation Figure 4The described device can perform a first characteristic superposition operation on each exposure linear data to obtain a first linear data corresponding to the exposure linear data, perform a second characteristic superposition operation on the first linear data corresponding to each exposure linear data to obtain a corresponding second linear data, and generate exposure output data according to the second linear data, wherein the first characteristic superposition operation includes camera response curve superposition operation and noise superposition operation; the second characteristic superposition operation includes motion blur superposition operation, inter-frame motion superposition operation and spatial misregistration superposition operation, and the exposure output data includes long-exposure pixel value data, medium-exposure pixel value data and short-exposure pixel value data corresponding to the exposure linear data. The pixel value can be adjusted by the first characteristic superposition operation to match the actual response characteristics of the camera and the noise superposition operation is performed on the exposure linear data to simulate and reduce image noise. The first linear data corresponding to each exposure linear data is subjected to motion blur superposition by the second characteristic superposition operation to simulate the blurring effect of moving objects on the exposure device, and inter-frame motion superposition is performed to handle the motion changes between consecutive frames, improve image stability in dynamic scenes, and spatial misregistration superposition is performed to adjust the pixel position to simulate actual motion and scene changes. By simulating the camera response and noise characteristics, the generated image is closer to the actual shooting effect. The motion blur and inter-frame motion superposition operations help better handle dynamic scenes and reduce motion artifacts. The spatial misregistration superposition operation can improve the stability and continuity of the image in fast motion or complex scenes. By processing long-exposure, medium-exposure and short-exposure pixel value data respectively, better exposure balance can be achieved in different lighting conditions, which can adapt to different shooting environments and lighting conditions and provide consistent image quality. Through detailed characteristic superposition operations, not only the quality and authenticity of the image are improved, but also the processing capability of dynamic scenes is enhanced, so that the final exposure output data can maintain high quality under various conditions. Through precise exposure control and data splitting, the processing effect of wide dynamic video RAW data is significantly improved, which is conducive to enhancing the image training data set and helping the deep learning network to improve the accuracy and intelligence of image processing in the actual multi-exposure wide dynamic image processing process, thereby improving the quality of image processing.
[0188] In yet another optional embodiment, as shown in FIG. 3B, the training data pair construction module 304 is configured to construct the training data pair based on all the exposure output data, and the specific manner of constructing the training data pair includes: Figure 4
[0189] performing a first ground truth data construction operation on the exposure output data based on the camera response curve, to obtain first ground truth construction data, wherein the first ground truth construction data comprises a long exposure pixel value of the exposure output data superimposed with the camera response curve, a medium exposure pixel value of the exposure output data superimposed with the camera response curve, a short exposure pixel value of the exposure output data superimposed with the camera response curve, and an ultra-short exposure pixel value of the exposure output data superimposed with the camera response curve;
[0190] performing a second ground truth data construction operation on the first ground truth construction data based on the noise processing parameter, to obtain second ground truth construction data, wherein the second ground truth construction data comprises a medium exposure pixel value of the first ground truth construction data superimposed with noise, a short exposure pixel value of the first ground truth construction data superimposed with noise, and an ultra-short exposure pixel value of the first ground truth construction data superimposed with noise;
[0191] generating target training data based on the first ground truth construction data and the second ground truth construction data, in combination with a pre-determined inter-frame motion superimposition ground truth data construction formula and a pre-determined spatial misregistration superimposition ground truth data construction formula;
[0192] obtaining a training data pair according to all the target training data.
[0193] It can be seen that, by implementing the method, the training data pair can be obtained, and the training data pair is used for training the neural network model, so that the neural network model can be trained to have a better performance. Figure 4The described device can perform a first ground truth data construction operation based on exposure output data and camera response curves to obtain first ground truth construction data, and perform a second ground truth data construction operation based on the first ground truth construction data and noise processing parameters to obtain second ground truth construction data, generate target training data based on the first ground truth data construction and the second ground truth data construction data combined with inter-frame motion superimposed ground truth data construction formula and spatial misregistration superimposed ground truth data construction formula to obtain training data pairs. The long exposure, medium exposure and short exposure pixel value data can be combined with the camera response curve based on the exposure output data and the camera response curve to simulate the real response when the camera captures the image, and the second ground truth construction data can be obtained based on the first ground truth construction data and the noise processing parameters. The simulation of the actual noise in the shooting can increase the complexity of the real world for the training data, and the generation of the target training data based on the first ground truth construction data, the second ground truth construction data, the inter-frame motion superimposed ground truth data construction formula and the spatial misregistration superimposed ground truth data construction formula can help to simulate the motion blur and spatial changes in the dynamic scene, and provide more comprehensive training information for the algorithm. Through the data construction containing inter-frame motion and spatial misregistration, the algorithm can more accurately process the dynamic scene, and through the superposition of noise in the training data, the algorithm can learn how to more effectively identify and reduce image noise. Through the training data pairs containing various complex situations, the adaptability of the model to different shooting conditions can be improved, the authenticity and diversity of the training data pairs can help to improve the accuracy of image analysis tasks (such as target detection and image segmentation), and through the construction of training data pairs containing various factors and complexity, the performance and robustness of the image processing algorithm can be significantly improved, especially in processing real-world image data. Through precise exposure control and data splitting, the processing effect of wide dynamic video RAW data can be significantly improved, which is beneficial to enhance the image training data set, and can help the deep learning network to improve the accuracy and intelligence of image processing in the actual multi-exposure wide dynamic image processing process, thereby improving the quality of image processing.
[0194] In yet another optional embodiment, as shown in FIG. 3B, the construction module 304 can construct the training data pairs based on all exposure output data in the following specific manner: Figure 4
[0195] Based on the exposure output data and the pre-determined inverse camera response curve, determine the wide dynamic linear ground truth data;
[0196] Based on the wide dynamic linear ground truth data, construct the training data pairs;
[0197] Furthermore, the conversion module 306 is also used to perform tone mapping operation on the wide dynamic range linear true value data according to the predetermined tone mapping algorithm after the construction module 304 determines the wide dynamic range linear true value data based on the exposure output data and the predetermined inverse camera response curve, so as to obtain tone mapping true value data.
[0198] Among them, the construction module 304, based on wide dynamic range linear ground truth data, constructs training data pairs by including the following specific methods:
[0199] Training data pairs are constructed based on ground truth tone mapping data.
[0200] It is evident that implementation Figure 4 The described apparatus is capable of determining wide dynamic range (WDR) linear ground truth data based on exposure output data and the inverse camera response curve, constructing training data pairs based on the WDR linear ground truth data, and further capable of performing tone mapping operations on the WDR linear ground truth data according to a tone mapping algorithm after determining the WDR linear ground truth data to obtain tone-mapped ground truth data, constructing training data pairs based on the tone-mapped ground truth data. It can utilize exposure output data and the inverse camera response curve to determine the WDR linear ground truth data, converting the data captured by the camera into linear space to better process wide dynamic range images, and using the determined WDR linear ground truth data to construct data pairs for training an image processing model. Furthermore, it can use the tone-mapped data to further construct training data pairs to train the model's color and brightness performance under different display conditions. Through the processing of the inverse camera response curve, details of high dynamic range scenes can be better recovered. Through the use of the tone mapping algorithm... The training data pairs, constructed to ensure color consistency and accuracy across different display devices, incorporate the complexity of wide dynamic range and color mapping, contributing to the training of more robust and accurate models. By including ground truth tone mapping data in the training process, tone mapping algorithms can be further optimized to adapt to different visual and display requirements. Through the construction of training data pairs, the model can learn how to present the best image effect on different display devices. This provides rich training information for image processing models by combining wide dynamic range linear ground truth data and tone mapping ground truth data, resulting in significant performance improvements in wide dynamic range imaging and color management. Consequently, it significantly enhances the processing effect of wide dynamic range video RAW data, which is beneficial for strengthening image training datasets and helping deep learning networks improve the accuracy and intelligence of image processing in actual multi-exposure wide dynamic range image processing, thereby improving the quality of image processing.
[0201] In yet another alternative embodiment, such as Figure 5As shown, the overlay module 303 performs a first feature overlay operation on each exposure linear data to obtain the first linear data corresponding to that exposure linear data in the following specific ways:
[0202] For each exposure linear data, a first feature superposition operation is performed on the exposure linear data according to a pre-determined camera response curve superposition algorithm and a pre-determined noise superposition algorithm to obtain the first linear data corresponding to the exposure linear data;
[0203] The pre-determined camera response curve overlay algorithm includes:
[0204] ;
[0205] in, These are the Bayer pixel values for long exposure, medium exposure, short exposure, and ultra-short exposure obtained from the preceding dynamic image segmentation. (.) includes the camera response curves under various illumination levels. These are the long exposure pixel values, medium exposure pixel values, short exposure pixel values, and ultra-short exposure pixel values after overlaying the camera response curve;
[0206] The predetermined noise superposition algorithm includes:
[0207] ;
[0208] in, These are the long exposure pixel values / medium exposure pixel values / short exposure pixel values / ultra-short exposure pixel values after overlaying the camera response curve. Including noise curves under various illuminance levels, These are the pixel values for long exposure, medium exposure, short exposure, and ultra-short exposure after adding noise.
[0209] Furthermore, for each exposure linear data point, a second feature overlay operation is performed on the first linear data corresponding to that exposure linear data point to obtain the second linear data corresponding to that exposure linear data point, including:
[0210] For each exposure linear data, according to the predetermined motion blur overlay algorithm, inter-frame motion overlay algorithm and spatial misalignment overlay algorithm, the second feature overlay operation is performed on the first linear data corresponding to the exposure linear data to obtain the second linear data corresponding to the exposure linear data.
[0211] The motion blur overlay algorithm includes:
[0212] ;
[0213] wherein, is the long / medium exposure image after motion blur superposition, is the long / medium exposure image after pre-sequential noise superposition, is a spatial blur kernel function that varies with exposure;
[0214] inter-frame motion superposition algorithm, comprising:
[0215] ;
[0216] wherein, , , represents the long / medium / short / ultra-short exposure image in time sequence, and its subscript represents the corresponding frame number on the time axis; , , represents the long / medium / short / ultra-short exposure image after reconstruction in time sequence;
[0217] spatial misregistration superposition algorithm, comprising:
[0218] ;
[0219] wherein, , , , represents the long / medium / short / ultra-short exposure image before sampling, and its subscript represents the spatial position coordinates of the corresponding pixel point, , , , represents the long / medium / short / ultra-short exposure image after spatial sampling reconstruction.
[0220] It can be seen that the implementation Figure 5The described device can combine the predetermined camera response curve superposition algorithm and the noise superposition algorithm to perform the first characteristic superposition operation on the first linear data corresponding to the exposure linear data, can generate an image closer to the real shooting effect by simulating the actual response and noise characteristics of the camera, and the superposition of the camera response curve helps to expand the dynamic range of the image, better preserve the details of the highlight and shadow parts, and the superposition of the noise model helps to consider the noise effect in the actual shooting in the image processing algorithm, thereby improving the effect of noise reduction processing, the training data pair contains camera response and noise information, which helps to improve the accuracy and reliability of the image processing algorithm in actual application, and the characteristic superposition operation can ensure the consistency and naturalness of the synthesis result when synthesizing multi-exposure high dynamic range images, and by considering the camera response and noise characteristics under different ISO settings, the method can adapt to different shooting environments and lighting conditions, and more real and accurate training data pairs in image analysis and computer vision tasks can improve the accuracy of task execution, and the automatic image processing process can utilize these characteristic superposition operations to reduce manual intervention and improve processing efficiency and consistency, thereby being able to provide more rich and real training data for image processing through the characteristic superposition operation by simulating the response and noise characteristics of the camera, thereby improving the performance and quality of image processing in multiple aspects, and thereby significantly improving the processing effect of wide dynamic video RAW data, which is beneficial to enhancing the image training data set, and helping the deep learning network to improve the accuracy and intelligence of image processing in the actual multi-exposure wide dynamic image processing process, thereby being beneficial to improving the quality of image processing.
[0221] Embodiment four
[0222] Please refer to Figure 5 , is another structure diagram of the training data construction device based on multi-exposure wide dynamic image disclosed by the embodiment of the application. As shown, the training data construction device based on multi-exposure wide dynamic image can include:
[0223] The memory 401 stores executable program codes;
[0224] The processor 402 is coupled to the memory 401;
[0225] The processor 402 calls the executable program codes stored in the memory 401 to execute the steps in the training data construction method based on multi-exposure wide dynamic image described in the embodiment one or the embodiment two of the application.
[0226] Embodiment five
[0227] The embodiment of the present application discloses a computer storage medium, which stores computer instructions, and the computer instructions are used to execute the steps of the training data construction method based on multi-exposure wide dynamic image described in the embodiment one or the embodiment two when being invoked.
[0228] Embodiment six
[0229] The embodiment of the present application discloses a computer program product, which comprises a non-transitory computer readable storage medium storing a computer program, and the computer program is operable to cause a computer to execute the steps of the training data construction method based on multi-exposure wide dynamic image described in the embodiment one or the embodiment two.
[0230] The above described device embodiments are only schematic, wherein the modules described as separated components can or can not be physically separated, and the components displayed as modules can or can not be physical modules, i.e., can be located in one place, or can be distributed on multiple network modules. Part or all of the modules can be selected according to actual needs to achieve the purpose of the present embodiment scheme. Those skilled in the art can understand and implement without creative labor.
[0231] Those skilled in the art can clearly understand the implementation of the various embodiments by means of software and necessary general hardware platforms through the above specific description of the embodiments, and of course, the embodiments can also be implemented by hardware. Based on such understanding, the above technical solutions, essentially or in the sense of contribution to the prior art, can be embodied in the form of a software product. The computer software product can be stored in a computer readable storage medium, which includes a Read-Only Memory (ROM), a Random Access Memory (RAM), a Programmable Read-only Memory (PROM), an Erasable Programmable Read Only Memory (EPROM), a One-time Programmable Read-Only Memory (OTPROM), an Electrically Erasable Programmable Read-Only Memory (EEPROM), a Compact Disc Read-Only Memory (CD-ROM), or other optical disk storage, magnetic disk storage, magnetic tape storage, or any other medium that can be used to carry or store data.
[0232] Finally, it should be noted that: the training data construction method and device based on multi-exposure wide dynamic image disclosed by the embodiments of the present application are only the preferred embodiments of the present application, and are used to illustrate the technical solutions of the present application, but not to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that; the technical solutions recorded in the foregoing embodiments can still be modified, or some technical features can be replaced by equivalents; and these modifications or replacements do not make the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present application.
Claims
1. A method for constructing training data based on multi-exposure wide dynamic range images, characterized in that, The method includes: Acquire target wide dynamic range video RAW data; Perform a data splitting operation on the target wide dynamic range video RAW data to obtain several exposure linear data; For each of the exposure linear data, a feature superposition operation is performed on the exposure linear data according to the predetermined feature superposition parameters to obtain exposure output data. The feature superposition parameters include one or more of the following: camera response curve superposition parameters, noise superposition parameters, motion blur superposition parameters, inter-frame motion superposition parameters, and spatial misalignment superposition parameters. Based on all the exposure output data, a training data pair is constructed; wherein, the training data pair includes the wide dynamic range mapped image corresponding to the target wide dynamic range video RAW data; After acquiring the target wide dynamic range video RAW data, and before performing a data splitting operation on the target wide dynamic range video RAW data to obtain several exposure linear data, the method further includes: Based on a predetermined pixel processing formula, pixel processing operations are performed on the target wide dynamic range video RAW data to obtain target linear data corresponding to the target wide dynamic range video RAW data, wherein the target linear data includes RGB data of the target wide dynamic range video RAW data after linearization processing; Perform a Bayer format conversion operation on the target linear data to obtain target format data, which includes any one of RG / GB, BG / GR, GR / BG, and GB / RG formats; The step of performing a data splitting operation on the target wide dynamic range video RAW data to obtain several exposure linear data points includes: Perform a data splitting operation on the target format data to obtain several exposure linear data; The construction of training data pairs based on all the exposure output data includes: Based on the exposure output data and the predetermined camera response curve, a first truth data construction operation is performed on the exposure output data to obtain first truth data construction data. The first truth data construction data includes long exposure pixel values after superimposing the exposure output data on the camera response curve, medium exposure pixel values after superimposing the exposure output data on the camera response curve, short exposure pixel values after superimposing the exposure output data on the camera response curve, and ultra-short exposure pixel values after superimposing the exposure output data on the camera response curve. Based on the first truth data and the predetermined noise processing parameters, a second truth data construction operation is performed on the first truth data to obtain the second truth data. The second truth data includes medium exposure pixel values after adding noise to the first truth data, short exposure pixel values after adding noise to the first truth data, and ultra-short exposure pixel values after adding noise to the first truth data. Based on the first and second ground truth construction data, and combined with the predetermined inter-frame motion superposition ground truth data construction formula and the predetermined spatial misalignment superposition ground truth data construction formula, target training data is generated. Based on all the target training data, training data pairs are obtained; The step of performing a first truth data construction operation on the exposure output data based on the exposure output data and a pre-determined camera response curve to obtain first truth data includes: Based on the exposure output data and the truth data construction algorithm corresponding to the pre-determined camera response curve, the first truth data construction operation is performed on the exposure output data to obtain the first truth data construction data. The algorithm for constructing the truth data corresponding to the camera response curve includes: ; in, The Bayer pixel values obtained when performing the data splitting operation on the target wide dynamic range video RAW data are: long exposure Bayer pixel values, medium exposure Bayer pixel values, short exposure Bayer pixel values, and ultra-short exposure Bayer pixel values. This represents the camera response curve that does not change with the camera's ISO. These are the long exposure pixel values, medium exposure pixel values, short exposure pixel values, and ultra-short exposure pixel values that do not change with camera ISO after being superimposed on the camera response curve; The noise processing parameters include: ; in, These are the medium exposure pixel values, short exposure pixel values, and ultra-short exposure pixel values that do not change with camera ISO after being overlaid with the camera response curve. The noise curve represents long exposure. These are the medium exposure pixel values, short exposure pixel values, and ultra-short exposure pixel values that do not change with the camera ISO after noise is added; For each of the exposure linear data points, a feature overlay operation is performed on the exposure linear data according to a predetermined feature overlay parameter to obtain exposure output data, including: For each of the exposure linear data, a first feature superposition operation is performed on the exposure linear data to obtain the first linear data corresponding to the exposure linear data; For each of the exposure linear data, a second feature superposition operation is performed on the first linear data corresponding to the exposure linear data to obtain the second linear data corresponding to the exposure linear data. For each of the exposure linear data, exposure output data is generated based on the second linear data corresponding to that exposure linear data; The first feature overlay operation includes camera response curve overlay operation and noise overlay operation; the second feature overlay operation includes motion blur overlay operation, inter-frame motion overlay operation, and spatial misalignment overlay operation; the exposure output data corresponding to each of the exposure linear data includes long exposure pixel value data, medium exposure pixel value data, short exposure pixel value data, and ultra-short exposure pixel value data corresponding to the exposure linear data. Wherein, for each of the exposure linear data, performing a first feature superposition operation on the exposure linear data to obtain the first linear data corresponding to the exposure linear data includes: For each exposure linear data, a first feature superposition operation is performed on the exposure linear data according to a pre-determined camera response curve superposition algorithm and a pre-determined noise superposition algorithm to obtain the first linear data corresponding to the exposure linear data; The pre-determined camera response curve overlay algorithm includes: ; in, The Bayer pixel values obtained when performing the data splitting operation on the target wide dynamic range video RAW data are: long exposure Bayer pixel values, medium exposure Bayer pixel values, short exposure Bayer pixel values, and ultra-short exposure Bayer pixel values. (.) includes the camera response curves under various illumination levels. This represents the long exposure pixel values, medium exposure pixel values, short exposure pixel values, and ultra-short exposure pixel values under various illumination conditions after overlaying the camera response curves. The predetermined noise superposition algorithm includes: ; in, These are the long exposure pixel values, medium exposure pixel values, short exposure pixel values, and ultra-short exposure pixel values under various illumination levels after the camera response curves are overlaid. Including noise curves under various illuminance levels, These are the pixel values for long exposure, medium exposure, short exposure, and ultra-short exposure at various illumination levels after adding noise. For each of the exposure linear data, performing a second feature superposition operation on the first linear data corresponding to that exposure linear data to obtain the second linear data corresponding to that exposure linear data includes: For each of the exposure linear data, according to the predetermined motion blur overlay algorithm, inter-frame motion overlay algorithm and spatial misalignment overlay algorithm, the second feature overlay operation is performed on the first linear data corresponding to the exposure linear data to obtain the second linear data corresponding to the exposure linear data.
2. The method for constructing training data based on multi-exposure wide dynamic range images according to claim 1, characterized in that, The data splitting operation performed on the target format data yields several exposure linear data points, including: Based on the target format data, determine the exposure data range corresponding to the target wide dynamic range video RAW data, and determine the exposure data relationship between the exposure data range and the actual dynamic range; Based on the exposure data relationship, determine the exposure ratio corresponding to the target format data; Based on the exposure ratio and the predetermined split pixel expression, a data splitting operation is performed on the target format data to obtain several exposure linear data.
3. The method for constructing training data based on multi-exposure wide dynamic range images according to claim 1, characterized in that, The construction of training data pairs based on all the exposure output data includes: Based on the exposure output data and the pre-determined inverse camera response curve, determine the wide dynamic range linear true value data; Based on the wide dynamic range linear truth data, training data pairs are constructed; Furthermore, after determining the wide dynamic range linear true value data based on the exposure output data and the pre-determined inverse camera response curve, the method further includes: According to a predetermined tone mapping algorithm, a tone mapping operation is performed on the wide dynamic range linear true value data to obtain tone mapping true value data; The construction of training data pairs based on the wide dynamic range linear ground truth data includes: Training data pairs are constructed based on the ground truth tone mapping data.
4. A training data construction device based on multi-exposure wide dynamic range images, characterized in that, The device includes: The acquisition module is used to acquire target wide dynamic range video RAW data; The splitting module is used to perform data splitting operations on the target wide dynamic range video RAW data to obtain several exposure linear data; The overlay module is used to perform a feature overlay operation on each of the exposure linear data according to a predetermined feature overlay parameter to obtain exposure output data. The feature overlay parameter includes one or more of the following: camera response curve overlay parameter, noise overlay parameter, motion blur overlay parameter, inter-frame motion overlay parameter, and spatial misalignment overlay parameter. A construction module is used to construct training data pairs based on all the exposure output data; wherein, the training data pairs include the wide dynamic range mapped image corresponding to the target wide dynamic range video RAW data; The device further includes: The processing module is configured to perform pixel processing on the target wide dynamic range video RAW data based on a predetermined pixel processing formula after the acquisition module acquires the target wide dynamic range video RAW data and before the splitting module performs data splitting operation on the target wide dynamic range video RAW data to obtain several exposure linear data, wherein the target linear data includes the RGB data of the target wide dynamic range video RAW data after linearization processing; The conversion module is used to perform Bayer format conversion on the target linear data to obtain target format data, wherein the target format data includes any one of RG / GB, BG / GR, GR / BG, and GB / RG formats. Specifically, the splitting module performs data splitting operations on the target wide dynamic range video RAW data to obtain several exposure linear data in the following ways: Perform a data splitting operation on the target format data to obtain several exposure linear data; The specific methods by which the construction module constructs training data pairs based on all the exposure output data include: Based on the exposure output data and the predetermined camera response curve, a first truth data construction operation is performed on the exposure output data to obtain first truth data construction data. The first truth data construction data includes long exposure pixel values after superimposing the exposure output data on the camera response curve, medium exposure pixel values after superimposing the exposure output data on the camera response curve, short exposure pixel values after superimposing the exposure output data on the camera response curve, and ultra-short exposure pixel values after superimposing the exposure output data on the camera response curve. Based on the first truth data and the predetermined noise processing parameters, a second truth data construction operation is performed on the first truth data to obtain the second truth data. The second truth data includes medium exposure pixel values after adding noise to the first truth data, short exposure pixel values after adding noise to the first truth data, and ultra-short exposure pixel values after adding noise to the first truth data. Based on the first and second ground truth construction data, and combined with the predetermined inter-frame motion superposition ground truth data construction formula and the predetermined spatial misalignment superposition ground truth data construction formula, target training data is generated. Based on all the target training data, training data pairs are obtained; The construction module performs a first truth data construction operation on the exposure output data based on the exposure output data and a pre-determined camera response curve. The specific methods for obtaining the first truth data construction include: Based on the exposure output data and the truth data construction algorithm corresponding to the pre-determined camera response curve, the first truth data construction operation is performed on the exposure output data to obtain the first truth data construction data. The algorithm for constructing the truth data corresponding to the camera response curve includes: ; in, The Bayer pixel values obtained when performing the data splitting operation on the target wide dynamic range video RAW data are: long exposure Bayer pixel values, medium exposure Bayer pixel values, short exposure Bayer pixel values, and ultra-short exposure Bayer pixel values. This represents the camera response curve that does not change with the camera's ISO. These are the long exposure pixel values, medium exposure pixel values, short exposure pixel values, and ultra-short exposure pixel values that do not change with camera ISO after being superimposed on the camera response curve; The noise processing parameters include: ; in, These are the medium exposure pixel values, short exposure pixel values, and ultra-short exposure pixel values that do not change with camera ISO after being overlaid with the camera response curve. The noise curve represents long exposure. These are the medium exposure pixel values, short exposure pixel values, and ultra-short exposure pixel values that do not change with the camera ISO after noise is added; For each exposure linear data point, the overlay module performs a feature overlay operation based on pre-determined feature overlay parameters to obtain exposure output data. The specific methods include: For each of the exposure linear data, a first feature superposition operation is performed on the exposure linear data to obtain the first linear data corresponding to the exposure linear data; For each of the exposure linear data, a second feature superposition operation is performed on the first linear data corresponding to the exposure linear data to obtain the second linear data corresponding to the exposure linear data. For each of the exposure linear data, exposure output data is generated based on the second linear data corresponding to that exposure linear data; The first feature overlay operation includes camera response curve overlay operation and noise overlay operation; the second feature overlay operation includes motion blur overlay operation, inter-frame motion overlay operation, and spatial misalignment overlay operation; the exposure output data corresponding to each of the exposure linear data includes long exposure pixel value data, medium exposure pixel value data, short exposure pixel value data, and ultra-short exposure pixel value data corresponding to the exposure linear data. Specifically, the overlay module performs a first feature overlay operation on each exposure linear data point to obtain the first linear data corresponding to that exposure linear data point. The specific method for this is as follows: For each exposure linear data, a first feature superposition operation is performed on the exposure linear data according to a pre-determined camera response curve superposition algorithm and a pre-determined noise superposition algorithm to obtain the first linear data corresponding to the exposure linear data; The pre-determined camera response curve overlay algorithm includes: ; in, The Bayer pixel values obtained when performing the data splitting operation on the target wide dynamic range video RAW data are: long exposure Bayer pixel values, medium exposure Bayer pixel values, short exposure Bayer pixel values, and ultra-short exposure Bayer pixel values. (.) includes the camera response curves under various illumination levels. This represents the long exposure pixel values, medium exposure pixel values, short exposure pixel values, and ultra-short exposure pixel values under various illumination conditions after overlaying the camera response curves. The predetermined noise superposition algorithm includes: ; in, These are the long exposure pixel values, medium exposure pixel values, short exposure pixel values, and ultra-short exposure pixel values under various illumination levels after the camera response curves are overlaid. Including noise curves under various illuminance levels, These are the pixel values for long exposure, medium exposure, short exposure, and ultra-short exposure at various illumination levels after adding noise. For each of the exposure linear data, performing a second feature superposition operation on the first linear data corresponding to that exposure linear data to obtain the second linear data corresponding to that exposure linear data includes: For each of the exposure linear data, according to the predetermined motion blur overlay algorithm, inter-frame motion overlay algorithm and spatial misalignment overlay algorithm, the second feature overlay operation is performed on the first linear data corresponding to the exposure linear data to obtain the second linear data corresponding to the exposure linear data.
5. A training data construction device based on multi-exposure wide dynamic range images, characterized in that, The device includes: Memory containing executable program code; A processor coupled to the memory; The processor calls the executable program code stored in the memory to execute the training data construction method based on multi-exposure wide dynamic range images as described in any one of claims 1-3.
6. A computer storage medium, characterized in that, The computer storage medium stores computer instructions, which, when invoked, are used to execute the training data construction method based on multi-exposure wide dynamic range images as described in any one of claims 1-3.
Citation Information
Patent Citations
HDR image reconstruction method based on Raw domain
CN116757959A