Image reconstruction method and device, equipment and storage medium
By fusing feature information from LDR image data and event stream data, and utilizing a high dynamic range output model, the image quality issues of high contrast and high-speed motion scenes in traditional video imaging technology are solved, achieving efficient and accurate HDR image reconstruction.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-30
- Publication Date
- 2026-03-17
AI Technical Summary
Traditional video imaging technology suffers from overexposure defects and image quality degradation in high-contrast and high-speed motion scenes. In particular, it is difficult to balance the exposure ratio and signal-to-noise ratio in high frame rate video processing, which affects the user's viewing experience.
By acquiring low dynamic range (LDR) image data and event stream data, feature fusion is performed, and a preset high dynamic range output model is used for decoding output to generate a high dynamic range (HDR) image.
It improves the accuracy and efficiency of image reconstruction, ensures clear details in both bright and dark areas, resolves motion blur and artifacts, and generates high-quality HDR images.
Smart Images

Figure CN121685347A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of image processing, in particular to an image reconstruction method and device, equipment and storage medium. BACKGROUND
[0002] Traditional video imaging technology is limited by the dynamic range and exposure time of the sensor, and there are region overexposure defects in high-contrast scenes (for example, indoor-outdoor transition areas under strong light) or high-speed motion scenes (for example, fast-moving objects in autonomous driving). High dynamic range (HDR) video reconstruction technology has become a widely used technology in the fields of photography, video games and high-end displays by expanding the limited brightness range in low dynamic range (LDR) images or videos.
[0003] Most HDR imaging technologies for traditional cameras rely on capturing and fusing multiple images with different exposure times. However, the optimal exposure ratio between LDR frame sequences with different exposure settings depends on the scene and is difficult to balance for various scenes captured in video over time, and the video frame rate is limited. In addition, autonomous driving and industrial detection often require processing high frame rate videos, but increasing the frame rate easily leads to a decrease in signal-to-noise ratio under low illumination, resulting in loss of image details and poor image quality, which seriously affects the user's viewing experience.
[0004] Therefore, how to accurately and efficiently complete HDR image reconstruction is a problem to be solved at present. SUMMARY
[0005] The main purpose of the present application is to provide an image reconstruction method, device, equipment and storage medium, which aims to improve the efficiency and accuracy of HDR image reconstruction.
[0006] In a first aspect, the present application provides an image reconstruction method, which comprises the following steps: obtaining low dynamic range (LDR) image data and event stream data, the LDR image data being obtained by an LDR shooting device shooting a target scene, and the event stream data being obtained by an event shooting device shooting the target scene; performing feature fusion on feature information in the LDR image data and feature information in the event stream data to obtain target feature data; decoding and outputting the target feature data through a preset high dynamic range (HDR) output model to obtain HDR image data of the target scene.
[0007] In a second aspect, the present application provides a video image reconstruction device, comprising an acquisition module, a feature fusion module and an output module, wherein: The acquisition module is configured to acquire low dynamic range (LDR) image data and event stream data, the LDR image data being obtained by an LDR imaging device from a target scene, and the event stream data being obtained by an event imaging device from the target scene. The feature fusion module is configured to perform feature fusion on feature information in the LDR image data and feature information in the event stream data to obtain target feature data. The output module is configured to decode and output the target feature data by using a preset high dynamic range (HDR) output model to obtain HDR image data of the target scene.
[0008] In a third aspect, the present application provides a computer device, comprising a processor, a memory, and a computer program stored in the memory and executable by the processor, wherein the computer program, when executed by the processor, implements the steps of the image reconstruction method described above.
[0009] In a fourth aspect, the present application provides a computer readable storage medium, wherein the computer readable storage medium stores a computer program, and the computer program, when executed by a processor, implements the steps of the image reconstruction method described above.
[0010] The present application provides an image reconstruction method, device, equipment and storage medium. The present application acquires low dynamic range (LDR) image data and event stream data, the LDR image data being obtained by an LDR imaging device from a target scene, and the event stream data being obtained by an event imaging device from the target scene. The present application performs feature fusion on feature information in the LDR image data and feature information in the event stream data to obtain target feature data. The present application decodes and outputs the target feature data by using a preset high dynamic range (HDR) output model to obtain HDR image data of the target scene. In the present application, feature fusion is performed on feature information in the LDR image data and the event stream data, and the target feature data is decoded and output by using a preset HDR output model, so that the HDR image data of the target scene can be accurately obtained. In the present application, the LDR image data and the event stream data are reconstructed into HDR images. The LDR image data ensures the details of the image in bright and dark areas, so that the image texture is clear. The event stream data can capture high-speed scene transient changes, solve motion blur and artifacts, and effectively generate high-quality HDR images. BRIEF DESCRIPTION OF DRAWINGS
[0011] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the drawings needed to be used in the embodiments description will be briefly introduced. Obviously, the drawings in the following description are some embodiments of the present application, and other drawings can be obtained by those skilled in the art without any creative effort on the basis of these drawings.
[0012] Figure 1 A flowchart of an image reconstruction method provided by the embodiments of the present application is shown in FIG. 1. Figure 2 A scene diagram of the image reconstruction method provided by the embodiments of the present application is shown in FIG. 2. Figure 3 A flowchart of a sub-step of the image reconstruction method in FIG. 1 is shown in FIG. 3. Figure 1 A flowchart of a sub-step of the image reconstruction method in FIG. 1 is shown in FIG. 3. Figure 4 A flowchart of another image reconstruction method provided by the embodiments of the present application is shown in FIG. 4. Figure 5 A schematic block diagram of a video image reconstruction device provided by the embodiments of the present application is shown in FIG. 5. Figure 6 A schematic block diagram of a sub-module of the video image reconstruction device in FIG. 5 is shown in FIG. 6. Figure 5 A schematic block diagram of a sub-module of the video image reconstruction device in FIG. 5 is shown in FIG. 6. Figure 7 A schematic block diagram of a computer device provided by the embodiments of the present application is shown in FIG. 7.
[0013] The implementation, functional features and advantages of the present application will be further described with reference to the embodiments and the drawings. DETAILED DESCRIPTION
[0014] The technical solutions in the embodiments of the present application will be described clearly and completely with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are some embodiments of the present application, but not all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without any creative effort belong to the protection scope of the present application.
[0015] The flowchart shown in the drawings is only an example, and it is not necessary to include all the contents and operations / steps, and it is not necessary to execute in the described order. For example, some operations / steps can be decomposed, combined or partially combined, and thus the actual execution order can be changed according to the actual situation.
[0016] Traditional video imaging technology is limited by the dynamic range and exposure time of the sensor, and has the problem of area overexposure in high-contrast scenes (for example, indoor-outdoor transition areas under strong light) or high-speed motion scenes (for example, fast-moving objects in autonomous driving). High dynamic range (HDR) video reconstruction technology, by extending the limited brightness range in low dynamic range (LDR) images or videos, has become a widely used technology in the fields of photography, video games, and high-end displays.
[0017] Most HDR imaging techniques for traditional cameras rely on capturing and fusing multiple images with different exposure times. However, the optimal exposure ratio between LDR frame sequences with different exposure settings depends on the scene and is difficult to balance over time for various scenes captured in videos with limited frame rates. In addition, autonomous driving and industrial detection often require processing high-frame-rate videos. However, increasing the frame rate easily leads to a decrease in signal-to-noise ratio under low illumination, resulting in loss of image details and poor image quality, which seriously affects the user's viewing experience.
[0018] To solve the above problems, embodiments of the present application provide an image reconstruction method, device, equipment and storage medium. The image reconstruction method comprises: acquiring low dynamic range (LDR) image data and event stream data, the LDR image data being obtained by an LDR shooting device shooting a target scene, and the event stream data being obtained by an event shooting device shooting the target scene; performing feature fusion on feature information in the LDR image data and feature information in the event stream data to obtain target feature data; and decoding and outputting the target feature data through a preset high dynamic range (HDR) output model to obtain HDR image data of the target scene.
[0019] The image reconstruction method can be applied in a computer device, which can be a mobile phone, a tablet computer, a notebook computer, a desktop computer, a personal digital assistant, a wearable device, or the like.
[0020] Some embodiments of the present application will be described in detail below with reference to the accompanying drawings. In the case of no conflict, the embodiments described below and the features in the embodiments can be combined with each other.
[0021] Please refer to Figure 1 , Figure 1 A flowchart of an image reconstruction method provided by an embodiment of the present application is shown.
[0022] As shown in Figure 1 , the image reconstruction method comprises steps S101 to S103.
[0023] In step S101, low dynamic range (LDR) image data and event stream data are acquired, the LDR image data being obtained by an LDR imaging device imaging a target scene, and the event stream data being obtained by an event imaging device imaging the target scene.
[0024] The target scene is a scene to be imaged, such as an outdoor building landscape scene, a high-speed moving car scene, a night star twinkling scene, etc. The LDR image data is obtained by an LDR imaging device imaging the target scene, and the event stream data is obtained by an event imaging device imaging the target scene. The sampling frequencies of the LDR imaging device and the event imaging device can be set to the same sampling frequency, so that the LDR image data and the event stream data are acquired at the same time, and the HDR image reconstruction is more accurate.
[0025] It should be noted that the LDR imaging device and the event imaging device can be selected according to actual conditions, and the embodiments of the present application do not make specific limitations thereon. For example, the LDR imaging device can be an LDR camera, and the event imaging device can be an event camera.
[0026] In some embodiments, the target scene is imaged by the LDR imaging device to obtain the LDR image data. The target scene is imaged by the event imaging device to obtain the event stream data. The LDR image data and the event stream data can be accurately acquired by the LDR imaging device and the event imaging device.
[0027] In some embodiments, the LDR imaging device and the event imaging device are integrated in a beam splitter optical system, and the beam splitter optical system is controlled to enable the LDR imaging device and the event imaging device to simultaneously image the target scene to obtain the LDR image data and the event stream data. The target scene is imaged by the beam splitter optical system, which can effectively ensure that the acquisition time points of the LDR image data and the event stream data coincide, thereby effectively improving the accuracy of image reconstruction.
[0028] It should be noted that the more coincident the acquisition time stamps of the LDR image data and the event stream data are, the higher the quality of the reconstructed HDR image is. Generally, the acquisition time error of the LDR image data and the event stream data is less than or equal to 10 microseconds.
[0029] As shown in FIG. 1, an LDR imaging device and an event imaging device are integrated in a beam splitter optical system, and the LDR imaging device and the event imaging device simultaneously image an outdoor night scene to obtain LDR image data and event stream data. Figure 2
[0030] In step S102, feature information in the LDR image data and feature information in the event stream data are fused to obtain target feature data.
[0031] The target feature data is obtained by fusing the feature information in the LDR image data and the feature information in the event stream data in a multi-modal manner.
[0032] In some embodiments, before the feature information in the LDR image data and the feature information in the event stream data are fused to obtain the target feature data, the method further includes: performing time alignment and data adjustment on the LDR image data and the event stream data based on a time axis to obtain LDR image data and event stream data of the same time window. By performing time alignment and data adjustment on the LDR image data and the event stream data, the accuracy of image reconstruction can be effectively improved.
[0033] In some embodiments, as shown in FIG. 1, step S102 includes sub-step S1021 to sub-step S1023. Figure 3
[0034] In sub-step S1021, feature information in the LDR image data is extracted to obtain LDR image feature data.
[0035] In some embodiments, the LDR image data is subjected to inverse camera response function (CRF) transformation to obtain LDR image features in a linear domain, and the LDR image features in the linear domain are subjected to logarithmic compression processing to obtain the LDR image feature data. By performing inverse CRF transformation and logarithmic compression processing on the LDR image data, the LDR image feature data can be accurately obtained.
[0036] In some embodiments, the LDR image data is subjected to inverse camera response function (CRF) transformation to obtain LDR image features in a linear domain, and the LDR image features in the linear domain are subjected to logarithmic compression processing to obtain the LDR image feature data. By performing inverse CRF transformation and logarithmic compression processing on the LDR image data, the LDR image feature data can be accurately obtained.
[0037] In some embodiments, the LDR image data is subjected to inverse camera response function (CRF) transformation to obtain LDR image features in a linear domain, and the LDR image features in the linear domain are subjected to logarithmic compression processing to obtain the LDR image feature data. By performing inverse CRF transformation and logarithmic compression processing on the LDR image data, the LDR image feature data can be accurately obtained. For the linear domain LDR image feature, the preset gain value and the linear domain LDR image feature are operated based on the preset logarithmic compression formula to obtain LDR image feature data. The preset gain value can be set according to actual conditions, and the embodiments of the present application do not make specific limitations thereto.
[0038] The sub-step S1022 is feature extraction on the feature information in the event stream data to obtain event feature data.
[0039] The event feature data is an event feature matrix, and the event feature matrix includes spatial position, luminance change parameter and time sequence of the event.
[0040] In some embodiments, the event stream data is segmented based on a target time window to obtain a plurality of continuous event segments, the target time window is determined based on the object movement speed in the target scene; and an event feature matrix is constructed for events in each event segment according to position coordinates and luminance information to obtain event feature data.
[0041] In some embodiments, the event stream data is segmented based on a target time window to obtain a plurality of continuous event segments, and the manner can be: obtaining the object movement speed in the target scene, the object movement speed being the speed of the object movement in the target scene collected by the event shooting device, obtaining a preset time window formula, the preset time window formula being , the target time window, the preset time length, the preset time coefficient, and the object movement speed in the target scene; the preset time length, the preset time coefficient and the object movement speed in the target scene are calculated based on the preset time window formula to obtain the target time window. The event stream data is segmented based on the target time window according to the time axis sequence to obtain a plurality of continuous event windows, and each event window in the target time window is taken as an event segment, wherein the event span of each event segment is the target time window, the preset time coefficient and the preset time length can be set according to actual conditions, and the embodiments of the present application do not make specific limitations thereto, for example, the preset time coefficient can be set to 0.5, and the preset time length can be set to 10 ms.
[0042] In some embodiments, the event feature matrix of the events in each event segment is constructed according to the position coordinates and the brightness information, and the event feature data can be obtained in the following manner: the coordinates (x, y) of each event in the event segment are labeled, the event brightness matrix is determined according to the brightness information change of the events at each position coordinate based on the time axis sequence, the event brightness matrix and the resolution are fused to obtain the event feature data, the event feature matrix includes the spatial position of the event, the brightness change parameter and the time sequence, for example, the event feature matrix is , is the event brightness matrix, the is the resolution of the event shooting device collecting the event stream data.
[0043] In some embodiments, the event brightness matrix can be determined in the following manner: a preset two-dimensional event matrix is constructed, the preset two-dimensional event matrix includes positive event elements and negative event elements, the brightness change and the change frequency of the events at each position coordinate on the time axis are counted, wherein the brightness change is the change of the current time point relative to the last time point; the change frequency of the brightness change is taken as a positive event element update parameter, and the change frequency of the brightness change is taken as a negative event element update parameter; the positive event element update parameters of each event are superimposed to obtain a total positive event element update parameter; the negative event element update parameters of each event are superimposed to obtain a total negative event element update parameter; the positive event elements in the preset two-dimensional event matrix are updated according to the total positive event element update parameter, and the negative event elements in the preset two-dimensional event matrix are updated according to the total negative event element update parameter, to obtain the event brightness matrix.
[0044] Sub-step S1023, the feature information of the LDR image feature data and the event feature data is fused to obtain the target feature data.
[0045] The double-branch feature fusion model is obtained, the double-branch feature fusion model includes an image feature extraction layer, an event feature extraction layer and a feature fusion layer, and the double-branch feature fusion model is obtained by pre-training a neural network model based on a plurality of sample data; the LDR image feature data is down-sampled and feature-extracted by the image feature extraction layer to obtain an LDR image feature map; the event feature data is dynamically feature-extracted by the event feature extraction layer to obtain an event feature map; and the LDR image feature map and the event feature map are feature-weighted and fused by the feature fusion layer to obtain target feature data. By weighting and fusing the feature information of the LDR image feature data and the event feature data, the target feature data can be accurately obtained. By weighting and fusing the feature information of the LDR image feature data and the event feature data by the double-branch feature fusion model, the target feature data can be accurately obtained.
[0046] It should be noted that the dual-branch feature fusion model is obtained by pre-training a neural network model based on a plurality of sample data, the sample data including sample LDR image feature data, sample event feature data, and labeled target feature data labels. The neural network model includes but is not limited to a convolutional neural network model and a recurrent convolutional neural network model, etc. The training process of the dual-branch feature fusion model can refer to the training process of the high dynamic range output model in the following embodiments.
[0047] In some embodiments, the manner of downsampling and feature extraction of the LDR image feature data by the image feature extraction layer to obtain the LDR image feature map can be: downsampling and feature extraction of the LDR image feature data by the image feature extraction layer to obtain a plurality of LDR image block features; and texture information feature extraction of each LDR image block feature to obtain an LDR image feature map, the LDR image feature map having clear texture details, making the image reconstruction more accurate.
[0048] In some embodiments, the manner of dynamic feature extraction of the event feature data by the event feature extraction layer to obtain the event feature map can be: dynamic feature extraction of the event feature data by the event feature extraction layer, the dynamic features including but not limited to motion trajectories and brightness change gradients; and capturing the sparsity and asynchrony of the event stream to obtain the event feature map.
[0049] In some embodiments, the manner of feature weighted fusion of the LDR image feature map and the event feature map by the feature fusion layer to obtain the target feature data can be: obtaining a first preset parameter matrix and a second preset parameter matrix; weighted calculation of the first preset parameter matrix and the LDR image feature map to obtain a target LDR image feature map; weighted calculation of the second preset parameter matrix and the event feature map to obtain a target event feature map; and feature convolution mixed fusion of the target LDR image feature map and the target event feature map to obtain the target feature data.
[0050] Step S103, decoding and outputting the target feature data by a preset high dynamic range output model to obtain high dynamic range HDR image data of the target scene.
[0051] The HDR image data can be an HDR video image or an HDR image.
[0052] In some embodiments, as shown in FIG. 2, Figure 4 The image reconstruction method further includes steps S201 to S205.
[0053] Step S201, obtaining a sample data set, the sample data set comprising a plurality of sample data, the sample data comprising sample target feature data and labeled HDR image data labels.
[0054] The sample data comprises sample target feature data and labeled HDR image data labels.
[0055] In some embodiments, historical target feature data and HDR image data corresponding to the historical target feature data are obtained; the historical target feature data is labeled as sample target feature data, and the HDR image data is labeled as HDR image data labels; the sample target feature data and the HDR image data labels are recorded as one sample data; and the foregoing steps are repeatedly executed to obtain the sample data set.
[0056] Step S202, obtaining a preset neural network model, and selecting one sample data from the sample data set as target sample data.
[0057] The preset neural network model comprises, but is not limited to, a convolutional neural network model and a recurrent convolutional neural network model. The preset neural network model comprises a convolutional layer and an output layer, the convolutional layer being used for up-sampling and convolutional processing of target feature data; and the output layer being used for normalization and decoding processing of HDR image data.
[0058] In some embodiments, one sample data is randomly selected from the sample data set as target sample data. The target sample data can be accurately obtained through the sample data set.
[0059] Step S203, decoding and outputting the sample target feature data in the target sample data through the preset neural network model to obtain predicted HDR image data.
[0060] The sample target feature data is up-sampled and convolutionally processed through the convolutional layer to obtain predicted candidate HDR image data; and the predicted candidate HDR image data is normalized and decoded through the output layer to obtain predicted HDR image data.
[0061] Step S204, determining whether the preset neural network model converges according to the predicted HDR image data and the HDR image data labels in the target sample data.
[0062] According to the predicted HDR image data and the HDR image data label in the target sample data, a loss value of the preset neural network model is determined; if the loss value is less than or equal to a preset loss value, it is determined that the preset neural network model has converged; if the loss value is greater than the preset loss value, it is determined that the preset neural network model has not converged. The preset loss value can be set according to actual conditions, and the embodiments of the present application do not make specific limitations thereto, for example, the preset loss value can be set to 0.02. By calculating the loss value of the preset neural network model, it can be accurately known whether the preset neural network model has converged.
[0063] In some embodiments, the manner of determining the loss value of the preset neural network model according to the predicted HDR image data and the HDR image data label in the target sample data can be: calculating the similarity of the predicted HDR image data and the HDR image data label to obtain a current similarity; obtaining a historical similarity, which is the average of the current similarities of the sample data that have completed training; performing mean calculation on the current similarity and the historical similarity to obtain a target similarity; subtracting the target similarity from 1 to obtain a value, which is determined as the model loss value. The manner of calculating the similarity can be selected according to actual conditions, and the embodiments of the present application do not make specific limitations thereto, for example, the cosine similarity of the predicted HDR image data and the HDR image data label is calculated.
[0064] In step S205, if the preset neural network model has not converged, the model parameters of the preset neural network model are adjusted, and the step of selecting a sample data from the sample data set as a target sample data is continued to be performed until a converged high dynamic range output model is obtained.
[0065] If the loss value of the preset neural network model is greater than the preset loss value, it is determined that the preset neural network model has not converged; the model parameters of the preset neural network model are adjusted, and the step of selecting a sample data from the sample data set as a target sample data is continued to be performed; the sample target feature data in the target sample data is decoded and output by the preset neural network model to obtain predicted HDR image data; according to the predicted HDR image data and the HDR image data label in the target sample data, it is determined whether the preset neural network model has converged; if the preset neural network model has not converged, the model parameters of the preset neural network model are adjusted until a converged high dynamic range output model is obtained.
[0066] In some embodiments, the target feature data is up-sampled and convoluted by a convolutional layer to obtain candidate HDR image data, and the candidate HDR image data is normalized and decoded by an output layer to obtain the HDR image data. The high dynamic range output model can accurately obtain the high dynamic range HDR image data of the target scene by decoding and outputting the target feature data, greatly improving the efficiency and accuracy of HDR image reconstruction.
[0067] The image reconstruction method provided by the above embodiments comprises the following steps: obtaining low dynamic range LDR image data and event stream data, the LDR image data being obtained by an LDR shooting device shooting a target scene, and the event stream data being obtained by an event shooting device shooting the target scene; performing feature fusion on feature information in the LDR image data and feature information in the event stream data to obtain target feature data; and decoding and outputting the target feature data by a preset high dynamic range output model to obtain high dynamic range HDR image data of the target scene. In this application, the feature information in the LDR image data and the event stream data is fused, and the target feature data is decoded and output by the preset high dynamic range output model, so that the HDR image data of the target scene can be accurately obtained. In this application, the LDR image data and the event stream data are reconstructed to obtain the HDR image data, the LDR image data ensures the details of the image in the highlight area and the weak light area, so that the image texture details are clear, the event stream data can capture the instantaneous changes of high-speed scenes, solve the motion blur and artifacts, and effectively generate high-quality HDR images.
[0068] Please refer to Figure 5 , Figure 5 A schematic block diagram of a video image reconstruction device provided by the embodiments of the present application is shown.
[0069] As Figure 5 shown, the video image reconstruction device 300 comprises an acquisition module 310, a feature fusion module 320 and an output module 330, wherein: The acquisition module 310 is configured to acquire low dynamic range LDR image data and event stream data, the LDR image data being obtained by an LDR shooting device shooting a target scene, and the event stream data being obtained by an event shooting device shooting the target scene; The feature fusion module 320 is configured to perform feature fusion on feature information in the LDR image data and feature information in the event stream data to obtain target feature data; The output module 330 is configured to decode and output the target feature data by a preset high dynamic range output model to obtain high dynamic range HDR image data of the target scene.
[0070] In some embodiments, as shown in FIG. 3, the feature fusion module 320 includes a feature extraction module 321 and a feature fusion sub-module 322, wherein: Figure 6 The feature extraction module 321 is configured to perform feature extraction on the feature information in the LDR image data to obtain LDR image feature data. The feature extraction module 321 is further configured to perform feature extraction on the feature information in the event stream data to obtain event feature data. The feature extraction module 321 is further configured to perform feature extraction on the feature information in the LDR image data to obtain LDR image feature data. The feature fusion sub-module 322 is configured to perform feature information fusion on the LDR image feature data and the event feature data to obtain the target feature data.
[0071] In some embodiments, the feature extraction module 321 is further configured to: perform inverse camera response function (CRF) transformation on the LDR image data to obtain linear domain LDR image features; perform logarithmic compression processing on the linear domain LDR image features to obtain the LDR image feature data.
[0072] In some embodiments, the feature extraction module 321 is further configured to: segment the event stream data based on a target time window to obtain a plurality of continuous event segments, the target time window being determined based on the object movement speed in the target scene; construct an event feature matrix for events in each of the event segments according to position coordinates and brightness information to obtain event feature data.
[0073] In some embodiments, the feature fusion sub-module 322 is further configured to: obtain a double-branch feature fusion model, the double-branch feature fusion model including an image feature extraction layer, an event feature extraction layer, and a feature fusion layer, the double-branch feature fusion model being obtained by pre-training a neural network model based on a plurality of sample data; perform down-sampling processing and feature extraction on the LDR image feature data through the image feature extraction layer to obtain an LDR image feature map; perform dynamic feature extraction on the event feature data through the event feature extraction layer to obtain an event feature map; perform feature weighted fusion on the LDR image feature map and the event feature map through the feature fusion layer to obtain the target feature data.
[0074] In some embodiments, the output module 330 is further configured to: perform up-sampling and convolution processing on the target feature data through the convolution layer to obtain candidate HDR image data. The candidate HDR image data is normalized and decoded by the output layer to obtain the HDR image data.
[0075] In some embodiments, the video image reconstruction apparatus 300 is further configured to: The LDR image data and the event stream data are aligned and adjusted based on a time axis to obtain LDR image data and event stream data of the same time window.
[0076] It should be noted that, for the convenience and brevity of description, the specific working process of the video image reconstruction apparatus can refer to the corresponding process in the foregoing image reconstruction method embodiments, which will not be described here.
[0077] Please refer to Figure 7 , Figure 7 The structural schematic block diagram of a computer device provided in the embodiments of the present application.
[0078] As Figure 7 shown, the computer device 400 includes a processor 402 and a memory 403 connected through a system bus 401, wherein the memory 403 can include a storage medium and an internal memory.
[0079] The storage medium can store a computer program. The computer program includes program instructions which, when executed, can cause the processor to execute any one of the image reconstruction methods.
[0080] The processor 402 is configured to provide computing and control capabilities to support the operation of the entire computer device.
[0081] The internal memory provides an environment for the execution of the computer program in the storage medium, and the computer program, when executed by the processor, can cause the processor to execute any one of the image reconstruction methods.
[0082] Those skilled in the art can understand that Figure 7 the structure shown in the foregoing embodiments is only a block diagram of part of the structure related to the scheme of the present application, and does not constitute a limitation on the computer device to which the scheme of the present application is applied. The specific computer device can include more or fewer components than those shown in the figure, or combine certain components, or have a different arrangement of components.
[0083] It should be appreciated that the processor 402 can be a central processing unit (CPU), the processor can also be other general-purpose processors, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA) or other programmable logic device, discrete gate or transistor logic device, discrete hardware component, etc. Among them, the general-purpose processor can be a microprocessor or the processor can also be any conventional processor, etc.
[0084] In one embodiment, the processor 402 is configured to run a computer program stored in the memory to implement the following steps: obtain low dynamic range (LDR) image data and event stream data, the LDR image data being obtained by an LDR imaging device from a target scene, and the event stream data being obtained by an event imaging device from the target scene; perform feature fusion on feature information in the LDR image data and feature information in the event stream data to obtain target feature data; decode and output the target feature data through a preset high dynamic range (HDR) output model to obtain HDR image data of the target scene.
[0085] In one embodiment, when implementing the feature fusion on the feature information in the LDR image data and the feature information in the event stream data to obtain the target feature data, the processor 402 is configured to implement: perform feature extraction on the feature information in the LDR image data to obtain LDR image feature data; perform feature extraction on the feature information in the event stream data to obtain event feature data; perform feature information fusion on the LDR image feature data and the event feature data to obtain the target feature data.
[0086] In one embodiment, when implementing the feature extraction on the feature information in the LDR image data to obtain the LDR image feature data, the processor 402 is configured to implement: perform inverse camera response function (CRF) transformation on the LDR image data to obtain LDR image features in a linear domain; perform logarithmic compression processing on the LDR image features in the linear domain to obtain the LDR image feature data.
[0087] In one embodiment, the processor 402, when implementing the feature extraction on the feature information in the event stream data to obtain event feature data, is configured to: segment the event stream data based on a target time window to obtain a plurality of continuous event segments, the target time window being determined based on the object movement speed in the target scene; construct an event feature matrix for events in each event segment according to position coordinates and brightness information to obtain event feature data.
[0088] In one embodiment, the processor 402, when implementing the feature information fusion on the LDR image feature data and the event feature data to obtain the target feature data, is configured to: obtain a double-branch feature fusion model, the double-branch feature fusion model including an image feature extraction layer, an event feature extraction layer, and a feature fusion layer, the double-branch feature fusion model being obtained by pre-training a neural network model based on a plurality of sample data; perform down-sampling processing and feature extraction on the LDR image feature data through the image feature extraction layer to obtain an LDR image feature map; perform dynamic feature extraction on the event feature data through the event feature extraction layer to obtain an event feature map; perform feature weighted fusion on the LDR image feature map and the event feature map through the feature fusion layer to obtain the target feature data.
[0089] In one embodiment, the preset high dynamic range output model includes a convolution layer and an output layer; the processor 402, when implementing the decoding output on the target feature data through the preset high dynamic range output model to obtain the high dynamic range (HDR) image data of the target scene, is configured to: perform up-sampling and convolution processing on the target feature data through the convolution layer to obtain candidate HDR image data; perform normalization and decoding processing on the candidate HDR image data through the output layer to obtain the HDR image data.
[0090] In one embodiment, before the processor implements the feature fusion on the feature information in the LDR image data and the feature information in the event stream data to obtain target feature data, the processor is further configured to: align the LDR image data and the event stream data based on a time axis to obtain LDR image data and event stream data of the same time window.
[0091] It should be noted that, for the convenience and brevity of description, the specific working process of the computer device is described above, and the corresponding process in the foregoing image reconstruction method embodiments can be referred to, which will not be described here.
[0092] The embodiments of the present application further provide a computer readable storage medium, which stores a computer program. The computer program includes program instructions. The method implemented by the program instructions can refer to the embodiments of the image reconstruction method of the present application.
[0093] The computer readable storage medium can be an internal storage unit of the computer device, such as a hard disk or a memory of the computer device. The computer readable storage medium can be non-volatile or volatile. The computer readable storage medium can also be an external storage device of the computer device, such as a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card, etc.
[0094] It should be understood that the terms used in the specification of the present application are only for the purpose of describing specific embodiments and are not intended to limit the present application. As used in the specification of the present application, unless the context clearly indicates otherwise, the singular forms "a", "an" and "the" are intended to include the plural forms.
[0095] It should also be understood that the term "and / or" used in the specification of the present application means any combination of one or more of the associated listed items and all possible combinations, and includes these combinations. It should be noted that in this document, the terms "include", "contain" or any other variants thereof are intended to cover non-exclusive inclusion, so that the process, method, article or system including a series of elements not only includes those elements, but also includes other elements not explicitly listed or inherent to such process, method, article or system. Without more limitations, the element defined by the statement "including a" does not exclude the presence of other identical elements in the process, method, article or system including the element.
[0096] The above-mentioned serial numbers of the embodiments of the present application are only for description, and do not represent the advantages and disadvantages of the embodiments. The above is only a specific implementation of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art can easily think of various equivalent modifications or replacements within the technical scope disclosed by the present application, and these modifications or replacements should be covered within the protection scope of the present application.
Claims
1. A method of image reconstruction, characterized by, The method comprises: acquiring low dynamic range (LDR) image data and event stream data, the LDR image data being obtained by an LDR camera shooting a target scene, and the event stream data being obtained by an event camera shooting the target scene; performing feature fusion on feature information in the LDR image data and feature information in the event stream data to obtain target feature data; decoding and outputting the target feature data through a preset high dynamic range (HDR) output model to obtain HDR image data of the target scene.
2. The image reconstruction method of claim 1, wherein, The feature fusion on the feature information in the LDR image data and the feature information in the event stream data to obtain the target feature data comprises: performing feature extraction on the feature information in the LDR image data to obtain LDR image feature data; performing feature extraction on the feature information in the event stream data to obtain event feature data; performing feature information fusion on the LDR image feature data and the event feature data to obtain the target feature data.
3. The image reconstruction method of claim 2, wherein, The feature extraction on the feature information in the LDR image data to obtain the LDR image feature data comprises: performing inverse camera response function (CRF) transformation on the LDR image data to obtain LDR image features in a linear domain; performing logarithmic compression processing on the LDR image features in the linear domain to obtain the LDR image feature data.
4. The image reconstruction method of claim 2, wherein, The feature extraction on the feature information in the event stream data to obtain the event feature data comprises: segmenting the event stream data based on a target time window to obtain a plurality of continuous event segments, the target time window being determined based on object movement speed in the target scene; constructing an event feature matrix according to position coordinates and brightness information of events in each event segment to obtain event feature data.
5. The image reconstruction method of claim 2, wherein, The feature information fusion on the LDR image feature data and the event feature data to obtain the target feature data comprises: acquiring a double-branch feature fusion model, the double-branch feature fusion model comprising an image feature extraction layer, an event feature extraction layer, and a feature fusion layer, the double-branch feature fusion model being obtained by pre-training a neural network model based on a plurality of sample data; performing down-sampling processing and feature extraction on the LDR image feature data through the image feature extraction layer to obtain an LDR image feature map; performing dynamic feature extraction on the event feature data through the event feature extraction layer to obtain an event feature map; performing feature weighted fusion on the LDR image feature map and the event feature map through the feature fusion layer to obtain the target feature data.
6. The image reconstruction method of claim 1, wherein, The preset HDR output model comprises a convolution layer and an output layer; and the decoding and outputting of the target feature data through the preset HDR output model to obtain the HDR image data of the target scene comprises: performing up-sampling and convolution processing on the target feature data through the convolution layer to obtain candidate HDR image data; The candidate HDR image data is normalized and decoded by the output layer to obtain the HDR image data.
7. The image reconstruction method of any one of claims 1-6, wherein, Before the feature fusion of the feature information in the LDR image data and the feature information in the event stream data to obtain target feature data, the method further includes: The LDR image data and the event stream data are aligned and adjusted in the same time by a time axis to obtain LDR image data and event stream data in the same time window.
8. A video image reconstruction apparatus characterized by comprising: The video image reconstruction device includes an acquisition module, a feature fusion module, and an output module, wherein: The acquisition module is configured to acquire low dynamic range (LDR) image data and event stream data, the LDR image data being obtained by an LDR shooting device shooting a target scene, and the event stream data being obtained by an event shooting device shooting the target scene; The feature fusion module is configured to perform feature fusion on feature information in the LDR image data and feature information in the event stream data to obtain target feature data; The output module is configured to decode and output the target feature data by a preset high dynamic range (HDR) output model to obtain HDR image data of the target scene.
9. A computer device, comprising: The computer device includes a processor, a memory, and a computer program stored on the memory and executable by the processor, wherein the computer program, when executed by the processor, implements the steps of the image reconstruction method according to any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that, The computer readable storage medium stores a computer program, wherein the computer program, when executed by a processor, implements the steps of the image reconstruction method according to any one of claims 1 to 7.
Citation Information
Patent Citations
Method and system for generating high dynamic range video under guidance of event camera
CN116456183A
Image processing method and related device thereof
CN119067879A
Tone mapping of high dynamic range images
US20070014470A1