HDR image generation method and device and electronic equipment

By acquiring the high and low-bit images of the LDR image and using the trained inverse tone mapping network model to perform inverse tone mapping, the problem of large image loss during HDR image generation in the prior art is solved, and a higher quality HDR image reconstruction is achieved.

CN120198339APending Publication Date: 2025-06-24JINAN BOGUAN INTELLIGENT TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311790668.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-12-22
Publication Date
2025-06-24

AI Technical Summary

Technical Problem

In the prior art, when generating HDR images using neural networks based on inverse tone mapping, there is a problem of large image loss, resulting in low quality of generated HDR images.

Method used

By acquiring the high- and low-bit images of the to-process LDR image, inversely mapping the high-bit image and the low-bit images based on the trained inverse tone mapping network model, the first map image corresponding to the high-bit image and the second map image corresponding to the low-bit image are obtained, and combined to generate the target HDR image.

Benefits of technology

This method can reconstruct the contour and high-frequency detail information of the HDR image more finely, reduce image artifacts and blurring, and improve the quality of the reconstructed HDR image.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120198339A_ABST
    Figure CN120198339A_ABST
Patent Text Reader

Abstract

The invention provides an HDR image generation method and device and electronic equipment, and relates to the technical field of image processing, and the method comprises the steps: obtaining a high-order image and a low-order image of a to-be-processed LDR image; performing inverse tone mapping on the high-order image and the low-order image based on an inverse tone mapping network model to obtain a first mapping image corresponding to the high-order image and a second mapping image corresponding to the low-order image; wherein the inverse tone mapping network model is obtained by training an initial inverse tone mapping network model based on a high-order sample image and a low-order sample image of a sample LDR image and an HDR label image corresponding to the sample LDR image; and combining the first mapping image and the second mapping image to generate a target HDR image corresponding to the LDR image to be processed. According to the technical scheme provided by the invention, the contour and high-frequency detail information of the HDR image can be reconstructed, image artifacts and blurring conditions are reduced, and the quality of the HDR image mapped by the LDR image is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image processing, and in particular, to an HDR image generation method, apparatus, and electronic device. Background Art

[0002] Compared with low dynamic range (LDR) images, high dynamic range (HDR) images can express clearer scene content in overexposed or underexposed scenes.

[0003] There are two ways to obtain HDR images. One is to directly obtain them using an HDR camera, and the other is to map LDR images into HDR images through an inverse tone mapping method. With the development of deep learning technology, in related technologies of inverse tone mapping, neural networks are gradually used for inverse tone mapping, and single or multiple LDR images with different exposure degrees are input into the inverse tone mapping neural network to reconstruct and generate HDR images. However, this method will cause image artifacts and blurring, resulting in significant image loss. Summary of the Invention

[0004] The present invention provides an HDR image generation method, apparatus, and electronic device to solve the problem of significant image loss when reconstructing and generating HDR images using a neural network based on inverse tone mapping in the prior art, so as to improve the quality of the reconstructed and generated HDR images.

[0005] The present invention provides an HDR image generation method, including:

[0006] Obtaining a high-bit image and a low-bit image of a to-be-processed LDR image;

[0007] Performing inverse tone mapping on the high-bit image and the low-bit image based on an inverse tone mapping network model to obtain a first mapped image corresponding to the high-bit image and a second mapped image corresponding to the low-bit image; wherein, the inverse tone mapping network model is obtained by training an initial inverse tone mapping network model based on a high-bit sample image and a low-bit sample image of a sample LDR image and an HDR label image corresponding to the sample LDR image;

[0008] Merging the first mapped image and the second mapped image to generate a target HDR image corresponding to the to-be-processed LDR image.

[0009] The present invention further provides an HDR image generation apparatus, including:

[0010] An obtaining module, configured to obtain a high-bit image and a low-bit image of a to-be-processed LDR image;

[0011] A mapping module, configured to perform inverse tone mapping on the high-bit image and the low-bit image based on an inverse tone mapping network model to obtain a first mapped image corresponding to the high-bit image and a second mapped image corresponding to the low-bit image; wherein, the inverse tone mapping network model is obtained by training an initial inverse tone mapping network model based on a high-bit sample image and a low-bit sample image of a sample LDR image and an HDR label image corresponding to the sample LDR image;

[0012] A generation module, configured to merge the first mapped image and the second mapped image to generate a target HDR image corresponding to the to-be-processed LDR image.

[0013] The present invention further provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor, wherein when the processor executes the computer program, the HDR image generation method according to any one of the above is implemented.

[0014] The present invention further provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the HDR image generation method according to any one of the above is implemented.

[0015] The present invention further provides a computer program product, including a computer program, and when the computer program is executed by a processor, the HDR image generation method according to any one of the above is implemented.

[0016] For the to-be-processed LDR image, the HDR image generation method, device and electronic device provided by the present invention first obtain the high-bit image and the low-bit image of the to-be-processed LDR image, and then perform inverse tone mapping on the high-bit image and the low-bit image based on the inverse tone mapping network model to obtain a first mapped image corresponding to the high-bit image and a second mapped image corresponding to the low-bit image. The inverse tone mapping network model is obtained by training an initial inverse tone mapping network model based on the high-bit sample image and the low-bit sample image of the sample LDR image and the HDR label image corresponding to the sample LDR image. Then, the first mapped image and the second mapped image are merged to generate a target HDR image corresponding to the to-be-processed LDR image, realizing the mapping from the LDR image to the HDR image. During the mapping from the LDR image to the HDR image, the high-bit image and the low-bit image of the LDR image are used. The high-bit image can reflect the contour information of the image, and the low-bit image can reflect the high-frequency detail information of the image. Compared with the method of inputting a single or multiple LDR images with different exposure degrees into an inverse tone mapping neural network to reconstruct and generate an HDR image, the contour and high-frequency detail information of the HDR image can be reconstructed more precisely, the situations of image artifacts and blurring are reduced, and the quality of the reconstructed and generated HDR image is improved. Description of the Drawings

[0017] To more clearly illustrate the technical solutions in the present invention or the prior art, the following will briefly introduce the accompanying drawings required for use in the description of the embodiments or the prior art. Obviously, the accompanying drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other accompanying drawings can also be obtained based on these drawings.

[0018] Figure 1 It is a schematic flowchart of the HDR image generation method provided by the embodiments of the present invention;

[0019] Figure 2 It is one of the schematic structural diagrams of the inverse tone mapping network model provided by the embodiments of the present invention;

[0020] Figure 3 It is a schematic flowchart of the construction method of the first query mapping table and the second query mapping table in the embodiments of the present invention;

[0021] Figure 4 It is another schematic structural diagram of the inverse tone mapping network model provided by the embodiments of the present invention;

[0022] Figure 5 It is yet another schematic structural diagram of the inverse tone mapping network model provided by the embodiments of the present invention;

[0023] Figure 6 It is a schematic structural diagram of the horizontal aggregation layer in the embodiments of the present invention;

[0024] Figure 7 It is a schematic structural diagram of the vertical aggregation layer in the embodiments of the present invention;

[0025] Figure 8 It is a schematic diagram of the principle of receptive field expansion in the embodiments of the present invention;

[0026] Figure 9 It is a schematic diagram of the principle of generating an HDR image based on the look-up table method in the embodiments of the present invention;

[0027] Figure 10 It is a schematic structural diagram of the HDR image generation device provided by the embodiments of the present invention;

[0028] Figure 11 It is a schematic structural diagram of the electronic device provided by the embodiments of the present invention. Detailed implementation manners

[0029] To make the objectives, technical solutions, and advantages of the present invention clearer, the technical solutions in the present invention will be clearly and completely described below with reference to the accompanying drawings in the present invention. Apparently, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present invention without creative efforts shall fall within the protection scope of the present invention.

[0030] It should be noted that the serial numbers assigned to the objects described in the present invention itself, such as "first", "second", etc., are only used to distinguish the described objects and do not have any sequential or technical meaning.

[0031] The following Figures 1-8 describes the HDR image generation method of the present invention. This HDR image generation method can be applied to electronic devices such as terminal devices or servers. Among them, terminal devices can include mobile phones, computers, in-vehicle devices, tablet computers, wearable devices, smart home devices, monitoring devices, etc.; servers can include independent servers, cluster servers, or cloud servers, etc. This HDR image generation method can also be applied to an HDR image generation device provided in an electronic device such as a terminal device or a server, and this HDR image generation device can be implemented by software, hardware, or a combination of both.

[0032] Figure 1 Exemplarily shows a schematic flowchart of the HDR image generation method provided by an embodiment of the present invention. Referring to Figure 1 as shown, this HDR image generation method can include the following steps 110 to 130.

[0033] Step 110: Obtain the high-bit image and the low-bit image of the LDR image to be processed.

[0034] The LDR image is an 8-bit precision image. Its first 4 bits can be taken to obtain the high-bit image, and its last 4 bits can be taken to obtain the low-bit image. That is, the high four-bit image (MSBs) and the low four-bit image (LSBs) of the LDR image are obtained.

[0035] Among them, the high four-bit image can reflect the contour information of the LDR image, and the low four-bit image can reflect the high-frequency detail information of the image.

[0036] Step 120: Perform inverse tone mapping on the high-bit image and the low-bit image based on the inverse tone mapping network model to obtain the first mapped image corresponding to the high-bit image and the second mapped image corresponding to the low-bit image.

[0037] Among them, the inverse tone mapping network model is obtained by training the initial inverse tone mapping network model based on the high-bit sample image, the low-bit sample image of the sample LDR image, and the HDR label image corresponding to the sample LDR image.

[0038] Specifically, a large number of sample LDR images and the HDR label images corresponding to the sample LDR images can be obtained, and the HDR label images are used as the label data during model training. After obtaining the sample LDR images, the high four-bit and low four-bit images of the sample LDR images can be extracted to obtain the corresponding high-bit sample images and low-bit sample images. Then, the high-bit sample images are input into the high-bit input channels of the initial inverse tone mapping network model, and the corresponding low-bit sample images are input into the low-bit input channels of the initial inverse tone mapping network model to obtain the high-bit prediction images and low-bit prediction images output by the initial inverse tone mapping network model. Then, the high-bit prediction images and low-bit prediction images are merged to obtain the HDR prediction images. Then, the loss function value is determined based on the HDR prediction images and the corresponding HDR label images, and the model parameters of the initial inverse tone mapping network model are adjusted using the loss function value until the model converges, obtaining the trained inverse tone mapping network model, which can map the LDR image to be processed to the HDR image.

[0039] Exemplarily, the inverse tone mapping network model may include a high-bit subnet and a low-bit subnet. The high-bit subnet is used to perform inverse tone mapping on the high-bit image of the LDR image to be processed and output the first mapped image corresponding to the high-bit image; the low-bit subnet is used to perform inverse tone mapping on the low-bit image of the LDR image to be processed and output the second mapped image corresponding to the low-bit image.

[0040] Exemplarily, the network structures of the high-bit subnet and the low-bit subnet are the same.

[0041] Step 130: Merge the first mapped image and the second mapped image to generate the target HDR image corresponding to the LDR image to be processed.

[0042] The first mapped image corresponding to the high-bit image is the HDR high-bit image after reconstruction of the high-bit image, and the second mapped image corresponding to the low-bit image is the HDR low-bit image after reconstruction of the low-bit image. After obtaining the first mapped image corresponding to the high-bit image and the second mapped image corresponding to the low-bit image, the target HDR image corresponding to the LDR image to be processed can be obtained by merging the first mapped image and the second mapped image.

[0043] The HDR image generation method provided by the embodiments of the present invention, for the to-be-processed LDR image, first obtains the high-bit image and the low-bit image of the to-be-processed LDR image, and then performs inverse tone mapping on the high-bit image and the low-bit image based on the inverse tone mapping network model to obtain the first mapped image corresponding to the high-bit image and the second mapped image corresponding to the low-bit image. Among them, the inverse tone mapping network model is obtained by training the initial inverse tone mapping network model based on the high-bit sample image and the low-bit sample image of the sample LDR image and the HDR label image corresponding to the sample LDR image. Then, the first mapped image and the second mapped image are merged to generate the target HDR image corresponding to the to-be-processed LDR image, realizing the mapping from the LDR image to the HDR image. In the process of mapping from the LDR image to the HDR image, the high-bit image and the low-bit image of the LDR image are used. The high-bit image can reflect the contour information of the image, and the low-bit image can reflect the high-frequency detail information of the image. Compared with the method of inputting a single or multiple LDR images with different exposure degrees into the inverse tone mapping neural network to reconstruct and generate the HDR image, it can reconstruct the contour and high-frequency detail information of the HDR image more precisely, reduce the situation of image artifacts and blurring, and improve the quality of the reconstructed HDR image.

[0044] Considering that it takes a long time for forward propagation when using a neural network for inverse tone mapping, and its time consumption is closely related to the structure of the neural network. That is, the deeper the neural network is longitudinally and the wider it is horizontally, the more complex the structure of the neural network is, and the higher the forward propagation time consumption is, but the better the inverse tone mapping effect is. In order to obtain a better inverse tone mapping effect based on the neural network in the mobile terminal with relatively weak inference performance or in the real-time video monitoring scenario, in an exemplary embodiment of the present invention, a look-up table method based on the neural network can be used to reconstruct the to-be-processed LDR image into the target HDR image. Specifically, all possible output results of the inverse tone mapping network model can be pre-stored in the query mapping table. In the actual inference stage, only the corresponding output data needs to be searched from the query mapping table to reconstruct the HDR image. In this way, not only can the characteristics of the neural network be utilized to achieve a good visual effect, but also the forward propagation of the traditional neural network is not required, greatly shortening the image reconstruction time, which is beneficial to the deployment and application of the inverse tone mapping method in the mobile terminal or real-time video monitoring scenario.

[0045] Based on this, combined with Figure 1 In an exemplary embodiment, performing inverse tone mapping on the high-bit image and the low-bit image based on the inverse tone mapping network model to obtain the first mapped image corresponding to the high-bit image and the second mapped image corresponding to the low-bit image may include:

[0046] Determine a first mapped image corresponding to the high-bit image and a second mapped image corresponding to the low-bit image by using a query mapping table corresponding to the inverse tone mapping network model; wherein, the query mapping table is used to store the correspondence between each preset query image block and the network mapping result; the size of the preset query image block is determined based on the receptive field of the spatial enhancement layer in the inverse tone mapping network model.

[0047] Wherein, the preset query image block is all possible image blocks formed by all pixel values used to represent the LDR image, and the size of the image block can be determined according to the receptive field of the spatial enhancement layer in the inverse tone mapping network model, such as being the same as the receptive field of the spatial enhancement layer in the inverse tone mapping network model. For example, if the receptive field of the spatial enhancement layer in the inverse tone mapping network model is 2×2, the size of the preset query image block can be 2×2, and the preset query image block is all combinations of pixel values with a size of 2×2. The spatial enhancement layer therein is used to perform spatial enhancement on the image input to the spatial enhancement layer to obtain a spatially enhanced feature map.

[0048] Exemplarily, each preset query image block can be input into the inverse tone mapping network model respectively to obtain the network mapping result output by the inverse tone mapping network model, and then the correspondence between each preset query image block and the network mapping result is generated into a query mapping table. Correspondingly, the query image blocks of the high-bit image can be traversed, and the network mapping result corresponding to each query image block can be found from the query mapping table to obtain the first mapped image corresponding to the high-bit image; similarly, the query image blocks of the low-bit image are traversed, and the network mapping result corresponding to each query image block is found from the query mapping table to obtain the second mapped image corresponding to the low-bit image. The query image block has the same size as the preset query image block.

[0049] Exemplarily, the inverse tone mapping network model may include a high-bit subnet and a low-bit subnet. The high-bit subnet is used to perform inverse tone mapping on the high-bit image of the LDR image to be processed and output a first mapped image corresponding to the high-bit image; the low-bit subnet is used to perform inverse tone mapping on the low-bit image of the LDR image to be processed and output a second mapped image corresponding to the low-bit image. The high-bit query image blocks in the preset query image blocks can be respectively input into the high-bit subnet to obtain the high-bit network mapping results output by the high-bit subnet, and then the corresponding relationship between the high-bit query image blocks and the high-bit network mapping results is determined as the high-bit query mapping table corresponding to the high-bit subnet; similarly, the low-bit query image blocks in the preset query image blocks can be respectively input into the low-bit subnet to obtain the low-bit network mapping results output by the low-bit subnet, and then the corresponding relationship between the low-bit query image blocks and the low-bit network mapping results is determined as the low-bit query mapping table corresponding to the low-bit subnet. Correspondingly, for the high-bit image of the LDR image to be processed, the query image blocks of the high-bit image can be traversed, and the network mapping results corresponding to each query image block can be found from the high-bit query mapping table, so as to obtain the first mapped image corresponding to the high-bit image; for the low-bit image of the LDR image to be processed, the query image blocks of the low-bit image can be traversed, and the network mapping results corresponding to each query image block can be found from the low-bit query mapping table, so as to obtain the second mapped image corresponding to the low-bit image. The query image blocks have the same size as the preset query image blocks.

[0050] In this way, by parallel table lookup, the efficiency of mapping the LDR image to be processed to the HDR image can be further improved.

[0051] Exemplarily, the inverse tone mapping network model can also be split according to the model structure, different query mapping tables are established for different modules, and the outputs of each module are determined by using the query mapping tables, which can reduce the forward propagation of the model and further shorten the time for HDR image reconstruction.

[0052] Based on this, in an exemplary embodiment of the present invention, Figure 2 Exemplarily shows one of the structural schematic diagrams of the inverse tone mapping network model provided by the embodiment of the present invention. Refer to Figure 2 As shown, the inverse tone mapping network model may include a high-bit subnet and a low-bit subnet, and the network structures of the high-bit subnet and the low-bit subnet are the same.

[0053] Among them, the high-bit subnet includes a high-bit spatial enhancement layer and a high-bit aggregation layer. The high-bit spatial enhancement layer is used to perform spatial enhancement on the high-bit image input into the high-bit subnet to obtain a first feature map after spatial enhancement; the high-bit aggregation layer is used to expand the receptive field of the first feature map obtained by the high-bit spatial enhancement layer to obtain a first expansion result. After the obtained first expansion result is fused with the high-bit image input into the high-bit subnet, a first mapped image corresponding to the high-bit image can be obtained.

[0054] The low - level sub - network includes a low - level spatial enhancement layer and a low - level aggregation layer. The low - level spatial enhancement layer is used to perform spatial enhancement on the low - level image input to the low - level sub - network to obtain a second feature map after spatial enhancement; the low - level aggregation layer is used to expand the receptive field of the second feature map obtained by the low - level spatial enhancement layer to obtain a second expansion result. After the obtained second expansion result is fused with the low - level image input to the low - level sub - network, a second mapped image corresponding to the low - level image can be obtained.

[0055] According to Figure 2 the inverse tone mapping network model shown, after obtaining the trained inverse tone mapping network model, the inverse tone mapping network model can be split to establish a first query mapping table corresponding to the high - level aggregation layer and a second query mapping table corresponding to the low - level aggregation layer, that is, the query mapping table can include a first query mapping table corresponding to the high - level aggregation layer of the high - level sub - network and a second query mapping table corresponding to the low - level aggregation layer of the low - level sub - network.

[0056] Specifically, Figure 3 an exemplary flow diagram of the construction method of the first query mapping table and the second query mapping table in the embodiments of the present invention is shown. Referring to Figure 3 shown, the first query mapping table and the second query mapping table are established based on the following steps 310 to 350.

[0057] Step 310: Obtain a preset query image block, which includes a high - level query image block and a low - level query image block.

[0058] Among them, the size of the preset query image block can be the same as the receptive field size of the spatial enhancement layer in the inverse tone mapping network model, that is, the same as the receptive fields of the high - level spatial enhancement layer and the low - level spatial enhancement layer. Among them, the high - level spatial enhancement layer and the low - level spatial enhancement layer have the same structure.

[0059] Specifically, the LDR image includes 256 pixel values, the low - level pixel values include 16 pixel values from 0 to 15, and the rest are high - level pixel values. Assuming that the receptive fields of the high - level spatial enhancement layer and the low - level spatial enhancement layer are 2×2, the low - level pixel values can be combined according to the size of 2×2 to obtain all the combination results, that is, all possible low - level query image blocks can be obtained. Similarly, the high - level pixel values can be combined according to the size of 2×2 to obtain all the combination results, that is, all possible high - level query image blocks can be obtained.

[0060] Step 320: Input each high - level query image block into the high - level aggregation layer respectively to obtain a first network mapping result corresponding to each high - level query image block output by the high - level aggregation layer.

[0061] Step 330: Determine the correspondence between the high-level query image blocks and the first network mapping result as the first query mapping table.

[0062] Step 340: Input each low-level query image block into the low-level aggregation layer respectively to obtain the second network output result corresponding to each low-level query image block output by the low-level aggregation layer.

[0063] Step 350: Determine the correspondence between the low-level query image blocks and the second network mapping result as the second query mapping table.

[0064] In this way, the first query mapping table corresponding to the high-level aggregation layer of the high-level sub-network and the second query mapping table corresponding to the low-level aggregation layer of the low-level sub-network can be established.

[0065] Combined Figure 2 and Figure 3 , in an exemplary embodiment of the present invention, using the query mapping table corresponding to the inverse tone mapping network model to determine the first mapped image corresponding to the high-level image and the second mapped image corresponding to the low-level image may include:

[0066] Input the high-level image into the high-level spatial enhancement layer of the high-level sub-network to obtain the first feature map output by the high-level spatial enhancement layer; use the first query mapping table to determine the first expansion result after the high-level aggregation layer expands the receptive field of the first feature map, and determine the first mapped image based on the first expansion result;

[0067] Input the low-level image into the low-level spatial enhancement layer of the low-level sub-network to obtain the second feature map output by the low-level spatial enhancement layer; use the second query mapping table to determine the second expansion result after the low-level aggregation layer expands the receptive field of the second feature map, and determine the second mapped image based on the second expansion result.

[0068] Among them, both the high-level spatial enhancement layer and the low-level spatial enhancement layer can be composed of two convolutional layers. The first convolution can use a 2×2 convolutional kernel to expand the spatial dimension to 64, and the second convolution can use a 1×1 convolutional kernel to increase the non-linear fitting ability of the network. The high-level spatial enhancement layer and the low-level spatial enhancement layer utilize the dependence of the input pixels in space and can obtain 16 feature maps of the input image.

[0069] The receptive field expansion may include receptive field expansion in the horizontal direction and receptive field expansion in the vertical direction.

[0070] Exemplarily, determining the first mapped image based on the first expansion result may include: fusing the first expansion result and the high-level image to obtain the first mapped image.

[0071] Exemplarily, determining the second mapped image based on the second expansion result may include: fusing the second expansion result and the low-level image to obtain the second mapped image.

[0072] In the process of using the look-up table method to determine the output result of a neural network, in order to reduce the memory occupation during table storage and the one-to-one correspondence between input and output, the look-up table method requires that the input search sequence be the same as the receptive field size of the network input. This results in a relatively low complexity of the neural network used to generate the query mapping table, a small receptive field, and insufficient correlation within the image neighborhood, making it difficult to obtain a good reconstruction effect. In order to enable the look-up table method to be better applied in complex HDR image reconstruction tasks and enhance the correlation within the image neighborhood to obtain an HDR image with better visual effects, in an exemplary embodiment of the present invention, the receptive field used by the look-up table method can be enlarged before look-up, decoupling the receptive field during look-up from the receptive field of the input image block, making it applicable to more complex HDR image reconstruction tasks.

[0073] Specifically, using the first query mapping table to determine the first expansion result after receptive field expansion of the high-level aggregation layer for the first feature map may include: performing pixel point misalignment fusion processing between at least two feature maps on the first feature map to obtain a third feature map; using the first query mapping table to determine the network mapping result corresponding to each query image block in the third feature map to obtain the target output result of the high-level aggregation layer; and determining the first expansion result after receptive field expansion of the high-level aggregation layer for the third feature map based on the target output result.

[0074] Wherein, the size of each query image block in the third feature map is the same as the size of the preset query image block in the query mapping table.

[0075] The high-level aggregation layer may include a horizontal aggregation layer and a vertical aggregation layer. The horizontal aggregation layer can be used to perform receptive field expansion in the horizontal direction, and the vertical aggregation layer can be used to perform receptive field expansion in the vertical direction. After obtaining the first feature map, the first feature map can be assigned to the horizontal aggregation layer and the vertical aggregation layer. A part of the first feature map is subjected to receptive field expansion in the horizontal direction, and another part of the first feature map is subjected to receptive field expansion in the vertical direction. Among them, the number of the horizontal aggregation layer and the vertical aggregation layer is equal.

[0076] Exemplarily, the number of input channels of the horizontal aggregation layer and the vertical aggregation layer is 2. The number of feature maps for a group of pixel point misalignment fusion can be determined according to the number of the first feature maps assigned to the horizontal aggregation layer and the vertical aggregation layer, so that there are two feature maps input to each horizontal aggregation layer and each vertical aggregation layer.

[0077] For example, assume that 16 first feature maps are obtained, and there are 2 horizontal aggregation layers and 2 vertical aggregation layers respectively. Then, each horizontal aggregation layer and each vertical aggregation layer can be assigned 4 first feature maps. For each horizontal aggregation layer or each vertical aggregation layer, the pixel points of these 4 first feature maps can be arranged in an interleaved manner between every 2 feature maps, so as to perform pixel misalignment addition on these 2 feature maps. After addition, the average value of the results is taken, and 2 third feature maps corresponding to the horizontal aggregation layer or the vertical aggregation layer can be obtained.

[0078] For another example, assume that 16 first feature maps are obtained, and there is 1 horizontal aggregation layer and 1 vertical aggregation layer respectively. Then, each horizontal aggregation layer and each vertical aggregation layer can be assigned 8 first feature maps. For each horizontal aggregation layer or each vertical aggregation layer, the pixel points of these 8 first feature maps can be arranged in a misaligned manner between every 4 feature maps, so as to perform sequential misalignment addition of pixel points on these 4 feature maps, that is, the second feature map is misaligned by one pixel point in the corresponding direction with respect to the first feature map, the third feature map is misaligned by one pixel point in the corresponding direction with respect to the second feature map, and so on. After addition, the average value of the results is taken, and 2 third feature maps corresponding to the horizontal aggregation layer or the vertical aggregation layer can be obtained.

[0079] Exemplarily, the target output result of the high-level aggregation layer includes the output results of each horizontal aggregation layer and each vertical aggregation layer. The average value of the output results of each layer can be taken to obtain the first expansion result after the high-level aggregation layer expands the receptive field of the third feature map.

[0080] In an exemplary embodiment of the present invention, the high-level aggregation layer can implement the expansion of the receptive field of the first feature map through one or more cascaded aggregation modules. Each aggregation module can include at least one horizontal aggregation layer and at least one vertical aggregation layer.

[0081] For example, in an alternative embodiment, in combination with Figure 2 , Figure 4 FIG. 2 exemplarily shows a second structural schematic diagram of the inverse tone mapping network model provided by the embodiment of the present invention. Referring to Figure 4 shown, in this inverse tone mapping network model, the high-level aggregation layer can include a first aggregation module. The first aggregation module includes at least one first horizontal aggregation layer and at least one first vertical aggregation layer, and the number of the first horizontal aggregation layer and the first vertical aggregation layer is the same. For example, in Figure 4 , 2 first horizontal aggregation layers and 2 first vertical aggregation layers are shown, namely the first horizontal aggregation layer A1, the first horizontal aggregation layer A2, the first vertical aggregation layer B1, and the first vertical aggregation layer B2. Among them, the first horizontal aggregation layer is used to expand the receptive field in the horizontal direction, and the first vertical aggregation layer is used to expand the receptive field in the vertical direction.

[0082] According to Figure 4The inverse tone mapping network model shown in the figure performs pixel dislocation fusion processing between at least two feature maps on the first feature map to obtain a third feature map, which may include:

[0083] The first feature map is allocated to each first horizontal aggregation layer and each first vertical aggregation layer; for each first horizontal aggregation layer, a pixel staggered fusion process is performed between at least two feature maps on the first feature map allocated to the first horizontal aggregation layer to obtain two third feature maps corresponding to the first horizontal aggregation layer; for each first vertical aggregation layer, a pixel staggered fusion process is performed between at least two feature maps on the first feature map allocated to the first vertical aggregation layer to obtain two third feature maps corresponding to the first vertical aggregation layer.

[0084] For example, assuming that 16 first feature maps are obtained, each first horizontal aggregation layer and each first vertical aggregation layer can be assigned 4 first feature maps respectively. Taking the first horizontal aggregation layer A1 as an example, the 4 first feature maps assigned to it can be divided into two groups, each with 2 feature maps. For each group, the pixels of the two feature maps in the group are displaced in the horizontal direction, such as one pixel is displaced, and the incomplete positions in the vertical direction can be filled with 0, and then the displaced pixels are merged and the average value is taken to obtain a third feature map corresponding to the two feature maps in the group. In this way, two third feature maps corresponding to the first horizontal aggregation layer A1 can be obtained. The same processing method can be used to obtain two third feature maps corresponding to each first horizontal aggregation layer and each first vertical aggregation layer.

[0085] In this way, the receptive field used by the lookup table method can be expanded through the staggered fusion of pixels.

[0086] In an optional embodiment, in combination Figure 4 , Figure 5 The third structural diagram of the inverse tone mapping network model provided by the embodiment of the present invention is exemplarily shown. Figure 5 As shown, in the inverse tone mapping network model, the high-order aggregation layer may further include at least one second aggregation module connected in series. Accordingly, the first query mapping table may include a first sub-query mapping table corresponding to the first aggregation module and second sub-query mapping tables corresponding to each second aggregation module. Figure 5 In the embodiment, the high-level aggregation layer may include a first aggregation module and a second aggregation module. When constructing the first query mapping table, a first sub-query mapping table corresponding to the first aggregation module and a second sub-query mapping table corresponding to the second aggregation module may be constructed.

[0087] Among them, the second aggregation module may include at least one second horizontal aggregation layer and at least one second vertical aggregation layer, and the number of the second horizontal aggregation layers is the same as that of the second vertical aggregation layers. The second horizontal aggregation layer is used to expand the receptive field in the horizontal direction, and the second vertical aggregation layer is used to expand the receptive field in the vertical direction. For example, in Figure 5 it exemplifies 2 second horizontal aggregation layers and 2 second vertical aggregation layers.

[0088] Specifically, in combination with Figure 3 and Figure 5 , step 320 may include: respectively inputting each high-level query image block into each first horizontal aggregation layer and each first vertical aggregation layer to obtain network mapping results respectively output by each first horizontal aggregation layer and each first vertical aggregation layer, and determining these network mapping results as the first sub-network mapping results corresponding to the high-level query image block in the first aggregation module; respectively inputting each high-level query image block into each second horizontal aggregation layer and each second vertical aggregation layer to obtain network mapping results respectively output by each second horizontal aggregation layer and each second vertical aggregation layer, and determining these network mapping results as the second sub-network mapping results corresponding to the high-level query image block in the second aggregation module; the first network mapping result includes the first sub-network mapping result corresponding to the first aggregation module and the second sub-network mapping result corresponding to the second aggregation module.

[0089] It can be understood that taking the first aggregation module including 4 aggregation layers as shown in Figure 4 as an example, for each high-level query image block, the first network mapping result corresponding to the high-level query image block may include network mapping results respectively corresponding to the 4 aggregation layers.

[0090] Exemplarily, a corresponding aggregation layer query mapping table may be established for each first horizontal aggregation layer and each first vertical aggregation layer respectively. Specifically, for each first horizontal aggregation layer, each high-level query image block may be respectively input into the first horizontal aggregation layer to obtain the aggregation layer network mapping result output by the first horizontal aggregation layer, and the corresponding relationship between the high-level query image block and the aggregation layer network mapping result is determined as the aggregation layer query mapping table corresponding to the first horizontal aggregation layer. Similarly, for each first vertical aggregation layer, each high-level query image block may be respectively input into the first vertical aggregation layer to obtain the aggregation layer network mapping result output by the first vertical aggregation layer, and the corresponding relationship between the high-level query image block and the aggregation layer network mapping result is determined as the aggregation layer query mapping table corresponding to the first vertical aggregation layer. The aggregation layer query mapping tables respectively corresponding to each first horizontal aggregation layer and each first vertical aggregation layer may constitute the first sub-query mapping table corresponding to the first aggregation module.

[0091] Similarly, corresponding aggregation layer query mapping tables can be established for each second horizontal aggregation layer and each second vertical aggregation layer respectively. The aggregation layer query mapping tables corresponding to each second horizontal aggregation layer and each second vertical aggregation layer can constitute the second sub-query mapping table corresponding to the second aggregation module.

[0092] Based on this, using the first query mapping table to determine the network mapping results corresponding to each query image block in the third feature map, and obtaining the target output result of the high-level aggregation layer, may include:

[0093] For each first horizontal aggregation layer, using the first sub-query mapping table to determine the network mapping results corresponding to each query image block in the two third feature maps corresponding to the first horizontal aggregation layer, and obtaining the first sub-output result of the first horizontal aggregation layer;

[0094] For each first vertical aggregation layer, using the first sub-query mapping table to determine the network mapping results corresponding to each query image block in the two third feature maps corresponding to the first vertical aggregation layer, and obtaining the second sub-output result of the first vertical aggregation layer;

[0095] Jointly determining the first sub-output results and the second sub-output results as the first output result of the first aggregation module, and determining the fourth feature map based on the first output result and the first feature map;

[0096] For each second aggregation module, using the second sub-query mapping table corresponding to the second aggregation module to determine the second output result after the receptive field of the fifth feature map is expanded by the second aggregation module; the fifth feature map is determined based on the output result of the previous aggregation module of the second aggregation module, and the input of the first second aggregation module is the fourth feature map;

[0097] Determining the second output result of the last second aggregation module as the target output result of the high-level aggregation layer.

[0098] Exemplarily, determining the fourth feature map based on the first output result and the first feature map includes: determining the average value of all the first sub-output results and all the second sub-output results in the first output result to obtain the first sub-expansion result corresponding to the first aggregation module; fusing the first sub-expansion result with each first feature map respectively to obtain the fourth feature map.

[0099] In this way, the number of fourth feature maps equal to the number of the first feature maps can be obtained.

[0100] Exemplarily, the fifth feature map is obtained by taking the average value of the output result of the previous aggregation module of the second aggregation module and then fusing it with each input feature map of the previous aggregation module respectively.

[0101] Exemplarily, determining the second output result after the receptive field of the fifth feature map is expanded by the second aggregation module using the second sub-query mapping table corresponding to the second aggregation module may include: performing pixel misalignment fusion processing between at least two feature maps on the fifth feature map to obtain a sixth feature map; using the second sub-query mapping table corresponding to the second aggregation module to determine the network mapping result corresponding to each query image block in the sixth feature map, so as to obtain the second output result after the receptive field of the fifth feature map is expanded by the second aggregation module.

[0102] Among them, performing pixel misalignment fusion processing between at least two feature maps on the fifth feature map may refer to the process of performing pixel misalignment fusion processing between at least two feature maps on the first feature map by the first aggregation module, which will not be elaborated here.

[0103] Exemplarily, determining the network mapping result corresponding to each query image block in the sixth feature map using the second sub-query mapping table to obtain the second output result after the receptive field of the fifth feature map is expanded by the second aggregation module may include: for each second horizontal aggregation layer, using the second sub-query mapping table to determine the network mapping result corresponding to each query image block in the two sixth feature maps corresponding to the second horizontal aggregation layer, so as to obtain the third sub-output result of the second horizontal aggregation layer; for each second vertical aggregation layer, using the second sub-query mapping table to determine the network mapping result corresponding to each query image block in the two sixth feature maps corresponding to the second vertical aggregation layer, so as to obtain the fourth sub-output result of the second vertical aggregation layer; jointly determining each third sub-output result and each fourth sub-output result as the second output result after the receptive field of the fifth feature map is expanded by the second aggregation module.

[0104] For the low-level aggregation layer, its network structure may be the same as that of the high-level aggregation layer, and a method similar to that of the high-level aggregation layer may be used to determine the second expansion result after the receptive field of the second feature map is expanded by the low-level aggregation layer.

[0105] Specifically, determining the second expansion result after the receptive field of the second feature map is expanded by the low-level aggregation layer using the second query mapping table may include: performing pixel misalignment fusion processing between at least two feature maps on the second feature map to obtain a seventh feature map; using the second query mapping table to determine the network mapping result corresponding to each query image block in the seventh feature map, so as to obtain the target output result of the low-level aggregation layer; determining the second expansion result after the receptive field of the seventh feature map is expanded by the low-level aggregation layer based on the target output result of the low-level aggregation layer. Among them, the size of each query image block in the seventh feature map is the same as the size of the preset query image block in the query mapping table.

[0106] Exemplarily, the low-level aggregation layer may include a horizontal aggregation layer and a vertical aggregation layer. The horizontal aggregation layer may be used to expand the receptive field in the horizontal direction, and the vertical aggregation layer may be used to expand the receptive field in the vertical direction. After obtaining the second feature map, the second feature map may be assigned to the horizontal aggregation layer and the vertical aggregation layer. A part of the second feature map is used to expand the receptive field in the horizontal direction, and another part of the second feature map is used to expand the receptive field in the vertical direction. Among them, the number of the horizontal aggregation layer and the vertical aggregation layer is equal.

[0107] Correspondingly, the target output result of the low-level aggregation layer includes the output results of each horizontal aggregation layer and each vertical aggregation layer. The average value of the output results of each layer may be taken to obtain the second expansion result after the low-level aggregation layer expands the receptive field of the seventh feature map.

[0108] In an exemplary embodiment of the present invention, the low-level aggregation layer may implement the expansion of the receptive field of the second feature map through one or more cascaded aggregation modules. Each aggregation module may include at least one horizontal aggregation layer and at least one vertical aggregation layer. Correspondingly, the second query mapping table may include sub-query mapping tables corresponding to each aggregation module.

[0109] Specifically, in an alternative embodiment, the low-level aggregation layer may include a third aggregation module. The third aggregation module includes at least one third horizontal aggregation layer and at least one third vertical aggregation layer, and the number of the third horizontal aggregation layer and the third vertical aggregation layer is the same. For example, referring to Figure 4 , the third aggregation module may include 2 third horizontal aggregation layers and 2 third vertical aggregation layers, namely, the third horizontal aggregation layer A3, the third horizontal aggregation layer A4, the third vertical aggregation layer B3, and the third vertical aggregation layer B4.

[0110] Exemplarily, a corresponding aggregation layer query mapping table may be established for each third horizontal aggregation layer and each third vertical aggregation layer. The aggregation layer query mapping tables corresponding to each third horizontal aggregation layer and each third vertical aggregation layer may constitute the third sub-query mapping table corresponding to the third aggregation module.

[0111] In an alternative embodiment, the low-level aggregation layer may further include at least one cascaded fourth aggregation module. Correspondingly, the second query mapping table includes the third sub-query mapping table corresponding to the third aggregation module and the fourth sub-query mapping tables corresponding to each fourth aggregation module. For example, referring to Figure 5 shown, the low-level aggregation layer may include a third aggregation module and a fourth aggregation module. When constructing the second query mapping table, the third sub-query mapping table corresponding to the third aggregation module and the fourth sub-query mapping table corresponding to the fourth aggregation module may be constructed.

[0112] Among them, the fourth aggregation module may include at least one fourth horizontal aggregation layer and at least one fourth vertical aggregation layer, and the number of the fourth horizontal aggregation layers is the same as that of the fourth vertical aggregation layers. For example, in Figure 5 two fourth horizontal aggregation layers and two fourth vertical aggregation layers are exemplified.

[0113] Specifically, in combination with Figure 3 and Figure 5 , step 340 may include: respectively inputting each low-order query image block into each third horizontal aggregation layer and each third vertical aggregation layer to obtain network mapping results respectively output by each third horizontal aggregation layer and each third vertical aggregation layer, and determining these network mapping results as the third sub-network mapping results corresponding to the low-order query image block corresponding to the third aggregation module; respectively inputting each low-order query image block into each fourth horizontal aggregation layer and each fourth vertical aggregation layer to obtain network mapping results respectively output by each fourth horizontal aggregation layer and each fourth vertical aggregation layer, and determining these network mapping results as the fourth sub-network mapping results corresponding to the low-order query image block corresponding to the fourth aggregation module; the second network mapping result includes the third sub-network mapping result corresponding to the third aggregation module and the fourth sub-network mapping result corresponding to the fourth aggregation module.

[0114] Exemplarily, a corresponding aggregation layer query mapping table may be established for each fourth horizontal aggregation layer and each fourth vertical aggregation layer respectively. The aggregation layer query mapping tables corresponding to each fourth horizontal aggregation layer and each fourth vertical aggregation layer may form the fourth sub-query mapping table corresponding to the fourth aggregation module.

[0115] Exemplarily, for the specific process of "using the second query mapping table to determine the network mapping result corresponding to each query image block in the seventh feature map to obtain the target output result of the low-order aggregation layer", reference may be made to the specific process of "using the first query mapping table to determine the network mapping result corresponding to each query image block in the third feature map to obtain the target output result of the high-order aggregation layer" in the embodiment corresponding to the high-order aggregation layer, which will not be elaborated here.

[0116] It can be understood that in the embodiments of the present invention, the inverse tone mapping network model may include a high-order sub-network for performing high-order image mapping and a low-order sub-network for performing low-order image mapping. The structures of the high-order sub-network and the low-order sub-network are the same, so the processing flows of the high-order sub-network and the low-order sub-network for the input images are the same. Only the high-order sub-network processes the high-order image of the to-be-processed LDR image, and the low-order sub-network processes the low-order image of the to-be-processed LDR image. And the query mapping tables corresponding to the two sub-networks are different, and the processing flows of the two sub-networks for their respective input images may refer to each other.

[0117] Based on the above embodiments, in combination with Figure 5The structure of the inverse tone mapping network model shown further illustrates the HDR image generation method provided by the embodiments of the present invention.

[0118] According to Figure 5 As shown, the inverse tone mapping network model may include two branches, a high-bit sub-network and a low-bit sub-network in parallel. The high-bit sub-network is used to process the high-bit image of the LDR image to be processed, and the low-bit sub-network is used to process the low-bit image of the LDR image to be processed. Finally, the processing results of the two branches are combined to obtain the HDR image corresponding to the LDR image to be processed.

[0119] Specifically, both the high-bit spatial enhancement layer in the high-bit sub-network and the low-bit spatial enhancement layer in the low-bit sub-network may be composed of two convolutional layers. The first convolution may use a 2×2 convolutional kernel to expand the spatial dimension to 64, and the second convolution may use a 1×1 convolutional kernel to increase the non-linear fitting ability of the network. The high-bit spatial enhancement layer and the low-bit spatial enhancement layer utilize the dependence of the input pixels in space and can obtain 16 feature maps of the input image.

[0120] The high-bit aggregation layer of the high-bit sub-network and the low-bit aggregation layer of the low-bit sub-network each include two aggregation modules, and the aggregation module can achieve the expansion of the receptive field. For each aggregation module, it may include at least one horizontal aggregation layer and at least one vertical aggregation layer. For example, in Figure 5 , taking each aggregation module including 2 horizontal aggregation layers and 2 vertical aggregation layers as an example, the horizontal aggregation layer can expand the receptive field size of the image width in the horizontal direction, and the vertical aggregation layer can expand the receptive field size of the image height in the vertical direction. The number of input channels for each horizontal aggregation layer and each vertical aggregation layer is 2.

[0121] In the high-bit spatial enhancement layer and the low-bit spatial enhancement layer, the pad set for the convolution is 0, and the length and width of the feature map obtained by the convolution will be reduced by 1. After the high-bit spatial enhancement layer outputs 16 first feature maps, every two first feature maps can be arranged with pixel misalignment, and the size of the fused image after the misalignment arrangement is kept the same as the size of the input image. For the incomplete pixel positions, 0 is filled. In this way, the receptive field can be expanded before the first feature map is input into the first aggregation module. The 16 second feature maps output by the low-bit spatial enhancement layer can be processed for receptive field expansion similar to that of the first feature map before being input into the third aggregation module.

[0122] For each horizontal aggregation layer, it may include 4 convolutional layers and 3 activation functions, and the activation function is, for example, the Gelu activation function. For example, Figure 6 Exemplarily shows the schematic structural diagram of the horizontal aggregation layer. Refer to Figure 6As shown, the first convolutional layer of the horizontal aggregation layer uses a 2×1 convolutional kernel, and the other three convolutional layers use 1×1 convolutional kernels. Through the horizontal aggregation layer, the receptive field of the image can be expanded in the horizontal direction.

[0123] For each vertical aggregation layer, it can include four convolutional layers and three activation functions, and the activation functions are, for example, Gelu activation functions. For example, Figure 7 An exemplary structural schematic diagram of the vertical aggregation layer is shown. Refer to Figure 7 As shown, the first convolutional layer of the vertical aggregation layer uses a 1×2 convolutional kernel, and the other three convolutional layers use 1×1 convolutional kernels. Through the vertical aggregation layer, the receptive field of the image can be expanded in the vertical direction.

[0124] In this way, through the 2×1 convolutional kernels of the horizontal aggregation layers and the 1×2 convolutional kernels of the vertical aggregation layers in the aggregation module, the receptive fields in the horizontal and vertical directions can be further expanded. For example, Figure 8 An exemplary schematic diagram of the principle of receptive field expansion is shown. Refer to Figure 8 As shown, first, the receptive fields of the high-level spatial enhancement layer and the low-level spatial enhancement layer are 2×2. After the pixel point misalignment fusion process, the receptive field is expanded to 3×3. After passing through the 2×1 convolutional kernels of the horizontal aggregation layers and the 1×2 convolutional kernels of the vertical aggregation layers in the first aggregation module, the receptive field is expanded to 16. After passing through the second aggregation module, the receptive field size can be expanded to 36. Skip connections are added between the aggregation modules to maintain the resulting features. In this way, when using the look-up table method to determine the network output, the receptive field of the network is larger, making the network deeper longitudinally, thereby improving the network performance and further making the visual quality of the finally obtained HDR image better.

[0125] According to Figure 5The inverse tone mapping network model shown can, during the model training phase, obtain the high-bit sample image and the low-bit sample image of the sample LDR image. The high-bit sample image is input into the high-bit spatial enhancement layer to obtain the first sample feature map output by the high-bit spatial enhancement layer. After pixel point misalignment arrangement and fusion of every two sample feature maps of the first sample feature map, the first fusion result is input into the first aggregation module to obtain the first prediction result output by each aggregation layer of the first aggregation module. After taking the average of each first prediction result, the average value is fused with each first sample feature map respectively through skip connection to obtain the second fusion result. Then the second fusion result is input into the second aggregation module to obtain the second prediction result output by each aggregation layer of the second aggregation module. Take the average of each second prediction result, and fuse this average value with the input high-bit sample image through skip connection to obtain the high-bit prediction image predicted by the high-bit sub-network. For the low-bit sample image, a processing flow similar to that of the high-bit sample image can be performed in the low-bit sub-network to obtain the low-bit prediction image predicted by the low-bit sub-network. Finally, the high-bit prediction image and the low-bit prediction image are merged to obtain the HDR prediction image corresponding to the sample LDR image. Then, based on this HDR prediction image and the HDR label image corresponding to the sample LDR image, the loss function value is determined, and the model parameters of the initial inverse tone mapping network model are adjusted using the loss function value until the model converges, and the trained inverse tone mapping network model can be obtained.

[0126] After training the inverse tone mapping network model, a query mapping table can be established for the model. When establishing the query mapping table, the high-bit sub-network and the low-bit sub-network can use the same operations. The networks of each branch can be separated, the high-bit spatial enhancement layer and the low-bit spatial enhancement layer are fixed, and the input of each aggregation module is constructed. After inputting into each aggregation module, the output of each aggregation module is obtained, and the corresponding relationship between the input and the output is generated into the query mapping table corresponding to each aggregation module.

[0127] Specifically, for the first aggregation module of each branch, an input sequence InBlock1[v1][v2][v3][v4] with a length of 4 can be constructed as the input of the first aggregation module and input into the 4 aggregation layers of the first aggregation module respectively. Among them, v1, v2, v3, and v4 are the pixel values of the preset query image blocks input, which can be used as the subscripts of the input sequence and represent a preset query image block. After inputting InBlock1[v1][v2][v3][v4] into the 4 aggregation layers of the first aggregation module, output sequences of the 4 aggregation layers can be obtained. The 4 aggregation layers can be sorted, and the output sequences can be represented as outBlock1_1[v1][v2][v3][v4], outBlock1_2[v1][v2][v3][v4], outBlock1_3[v1][v2][v3][v4], and outBlock1_4[v1][v2][v3][v4] in order. According to the correspondence between the input sequence and the output sequence, the aggregation layer query mapping table corresponding to each aggregation layer in the first aggregation module can be obtained, and each aggregation layer query mapping table can form the sub-query mapping table corresponding to the first aggregation module. In this way, the first sub-query mapping table corresponding to the first aggregation module in the high-order sub-network and the third sub-query mapping table corresponding to the third aggregation module in the low-order sub-network can be obtained.

[0128] Similarly, for the second aggregation module of each branch, an input sequence InBlock2[v1][v2][v3][v4] of the second aggregation module can be constructed and input into the 4 aggregation layers of the second aggregation module respectively. Output sequences of the 4 aggregation layers in the second aggregation module can be obtained. The 4 aggregation layers can be sorted, and the output sequences can be represented as outBlock2_1[v1][v2][v3][v4], outBlock2_2[v1][v2][v3][v4], outBlock2_3[v1][v2][v3][v4], and outBlock2_4[v1][v2][v3][v4] in order. According to the correspondence between the input sequence and the output sequence, the aggregation layer query mapping table corresponding to each aggregation layer in the second aggregation module can be obtained, and each aggregation layer query mapping table can form the sub-query mapping table corresponding to the second aggregation module. In this way, the second sub-query mapping table corresponding to the second aggregation module in the high-order sub-network and the fourth sub-query mapping table corresponding to the fourth aggregation module in the low-order sub-network can be obtained.

[0129] Exemplarily, for the aggregation module in the high-order sub-network branch, an input sequence can be constructed using high-order pixel values. For the aggregation module in the low-order sub-network branch, an input sequence can be constructed using low-order pixel values.

[0130] Exemplarily, when constructing the query mapping table of the aggregation module, for the high-order sub-network branch, the high-order pixel values can be quantized. For example, 240 pixel values of the high four bits can be quantized into 60. In this way, all possible pixel combinations of the high-order pixel values are 60×60×60×60. That is, in the high-order sub-network branch, the size of the query mapping table of each aggregation layer in each aggregation module is [60 4 , 1]. In this way, the efficiency of table lookup can be improved, and the memory occupancy can also be reduced.

[0131] For the low-order sub-network branch, since it is a low four-bit image and there are only 16 pixel values, quantization may not be necessary. The size of the query mapping table of each aggregation layer in each aggregation module is [16 4 , 1].

[0132] After constructing the query mapping table of the inverse tone mapping network model, the query mapping table can be saved. In the process of using the inverse tone mapping network model to map the LDR image to be processed into an HDR image, the query mapping table can be used to find the network mapping results corresponding to each query image block, so as to realize the fast mapping of the HDR image.

[0133] Specifically, Figure 9 Exemplarily shows the schematic diagram of the principle of generating an HDR image based on the table lookup method in the embodiment of the present invention. Referring to Figure 9 as shown, for the LDR image to be processed, the high-order image and the low-order image of the LDR image to be processed can be obtained first. The high-order image is input into the high-order spatial enhancement layer to obtain 16 first feature maps output by the high-order spatial enhancement layer; the first feature maps are subjected to pixel point misalignment fusion processing between two feature maps to obtain the third feature map; each 2×2-sized query image block in the third feature map is traversed, and the network mapping results corresponding to each query image block are determined by using the first sub-query mapping table corresponding to the first aggregation module, so as to obtain the outputs of 4 aggregation layers in the first aggregation module. Then, the average value of the outputs of these 4 aggregation layers is taken and fused with the first feature map to obtain 16 new feature maps. The 16 feature maps are again subjected to pixel point misalignment fusion processing between two feature maps. Similarly, each query image block in the obtained feature maps is traversed, and the outputs of 4 aggregation layers in the second aggregation module are obtained by querying the second sub-query mapping table corresponding to the second aggregation module. The average value of the outputs of these 4 aggregation layers is taken and fused with the input high-order image, and the first mapped image corresponding to the high-order image can be obtained. Among them, in order to facilitate the representation of the aggregation layer query mapping table corresponding to each aggregation layer, in Figure 9 , the label of the aggregation layer query mapping table is taken to be the same as the label of its corresponding aggregation layer.

[0134] For the low-level image of the LDR image to be processed, similar processing to that of the high-level image can be performed in the low-level sub-network to obtain a second mapped image corresponding to the low-level image.

[0135] After that, the obtained first mapped image and the second mapped image are merged to obtain the target HDR image corresponding to the LDR image to be processed. In this way, the reconstruction of the LDR image to the HDR image is realized through the neural network model based on the parallel look-up table method, the network inference process in the inverse tone mapping of the traditional neural network is removed, and the speed of the inverse tone mapping method of the neural network is greatly improved. At the same time, the characteristics of strong generalization and high fitting degree for non-linear functions of the neural network can be utilized to obtain an HDR image with better visual effects.

[0136] Exemplarily, during the look-up table process, skip connections can be used between adjacent two aggregation modules to maintain the invariance of image features.

[0137] The HDR image generation method provided by the embodiments of the present invention, on the one hand, utilizes the high-level image and the low-level image of the LDR image to reconstruct the HDR image. Compared with the method of inputting a single or multiple LDR images with different exposure degrees into the inverse tone mapping neural network to reconstruct and generate the HDR image, it can reconstruct the contour and high-frequency detail information of the HDR image more precisely, reduce the image artifacts and blurring conditions, and improve the visual effect and quality of the reconstructed and generated HDR image; on the other hand, during the HDR image reconstruction process, the pixel point misalignment fusion method between feature maps is used to expand the receptive field during the look-up table, decouple the size of the receptive field during the look-up table from the size of the input image block, and use a larger receptive field to enhance the correlation of the neural network neighborhood and the fitting degree of the neural network, thereby improving the visual effect of the HDR image; furthermore, the method of using multiple branches and serial look-up tables is used to reconstruct the HDR image, and the aggregation layer with the ability to expand the receptive field is created as the look-up mapping table, increasing the receptive field used by the look-up table method, enabling it to be applicable to more complex HDR image reconstruction tasks, and no interpolation operation is required after the look-up table, reducing the time complexity of HDR image generation and improving the efficiency of HDR image generation, which is beneficial to the deployment and application on simple devices such as mobile terminals.

[0138] Next, the HDR image generation device provided by the present invention will be described. The HDR image generation device described below can be mutually referred to the HDR image generation method described above.

[0139] Figure 10 Exemplarily shows the structural schematic diagram of the HDR image generation device provided by the embodiments of the present invention. Refer to Figure 10As shown in the figure, the HDR image generation device may include: an acquisition module 1010, configured to acquire a high-bit image and a low-bit image of the LDR image to be processed; a mapping module 1020, configured to perform inverse tone mapping on the high-bit image and the low-bit image based on an inverse tone mapping network model to obtain a first mapped image corresponding to the high-bit image and a second mapped image corresponding to the low-bit image, where the inverse tone mapping network model is obtained by training an initial inverse tone mapping network model based on a high-bit sample image and a low-bit sample image of a sample LDR image and an HDR label image corresponding to the sample LDR image; and a generation module 1030, configured to merge the first mapped image and the second mapped image to generate a target HDR image corresponding to the LDR image to be processed.

[0140] In one exemplary embodiment, the mapping module 1020 may be specifically configured to: determine a first mapped image corresponding to the high-bit image and a second mapped image corresponding to the low-bit image by using a query mapping table corresponding to the inverse tone mapping network model; where the query mapping table is used to store the correspondence between each preset query image block and the network mapping result; and the size of the preset query image block is determined based on the receptive field of the spatial enhancement layer in the inverse tone mapping network model.

[0141] In one exemplary embodiment, the inverse tone mapping network model includes a high-bit sub-network and a low-bit sub-network, and the query mapping table includes a first query mapping table corresponding to the high-bit aggregation layer of the high-bit sub-network and a second query mapping table corresponding to the low-bit aggregation layer of the low-bit sub-network. Correspondingly, the mapping module 1020 may include: a first enhancement unit, configured to input the high-bit image into the high-bit spatial enhancement layer of the high-bit sub-network to obtain a first feature map output by the high-bit spatial enhancement layer; a first determination unit, configured to determine, by using the first query mapping table, a first expansion result of the high-bit aggregation layer after expanding the receptive field of the first feature map, and determine the first mapped image based on the first expansion result; a second enhancement unit, configured to input the low-bit image into the low-bit spatial enhancement layer of the low-bit sub-network to obtain a second feature map output by the low-bit spatial enhancement layer; and a second determination unit, configured to determine, by using the second query mapping table, a second expansion result of the low-bit aggregation layer after expanding the receptive field of the second feature map, and determine the second mapped image based on the second expansion result.

[0142] In one exemplary embodiment, the first determination unit includes: a misalignment fusion subunit, configured to perform pixel misalignment fusion processing between at least two feature maps of the first feature map to obtain a third feature map; a first determination subunit, configured to determine, by using the first query mapping table, the network mapping result corresponding to each query image block in the third feature map to obtain a target output result of the high-bit aggregation layer; and a second determination subunit, configured to determine a first expansion result of the high-bit aggregation layer after expanding the receptive field of the third feature map based on the target output result.

[0143] In an exemplary embodiment, the high-level aggregation layer includes a first aggregation module, and the first aggregation module includes at least one first horizontal aggregation layer and at least one first vertical aggregation layer, and the number of the first horizontal aggregation layers is the same as that of the first vertical aggregation layers. Correspondingly, the dislocation fusion subunit is specifically configured to: allocate the first feature map to each first horizontal aggregation layer and each first vertical aggregation layer; for each first horizontal aggregation layer, perform pixel dislocation fusion processing between at least two feature maps of the first feature map allocated to the first horizontal aggregation layer to obtain two third feature maps corresponding to the first horizontal aggregation layer; for each first vertical aggregation layer, perform pixel dislocation fusion processing between at least two feature maps of the first feature map allocated to the first vertical aggregation layer to obtain two third feature maps corresponding to the first vertical aggregation layer.

[0144] In an exemplary embodiment, the high-level aggregation layer further includes at least one second aggregation module connected in series, and the first query mapping table includes a first sub-query mapping table corresponding to the first aggregation module and second sub-query mapping tables corresponding to the second aggregation modules respectively; Correspondingly, the first determination subunit is specifically configured to: for each first horizontal aggregation layer, use the first sub-query mapping table to determine the network mapping result corresponding to each query image block in the two third feature maps corresponding to the first horizontal aggregation layer to obtain a first sub-output result of the first horizontal aggregation layer; for each first vertical aggregation layer, use the first sub-query mapping table to determine the network mapping result corresponding to each query image block in the two third feature maps corresponding to the first vertical aggregation layer to obtain a second sub-output result of the first vertical aggregation layer; jointly determine each first sub-output result and each second sub-output result as a first output result of the first aggregation module, and determine a fourth feature map based on the first output result and the first feature map; for each second aggregation module, use the second sub-query mapping table corresponding to the second aggregation module to determine a second output result of the second aggregation module after expanding the receptive field of the fifth feature map; the fifth feature map is determined based on the output result of the previous aggregation module of the second aggregation module, and the input of the first second aggregation module is the fourth feature map; determine the second output result of the last second aggregation module as the target output result of the high-level aggregation layer.

[0145] In an exemplary embodiment, when the first determination unit determines the first mapped image based on the first expansion result, it is specifically configured to: fuse the first expansion result and the high-level image to obtain the first mapped image.

[0146] In an exemplary embodiment, the first query mapping table and the second query mapping table are established based on the following steps: obtaining preset query image blocks, where the preset query image blocks include high-level query image blocks and low-level query image blocks; inputting each high-level query image block into a high-level aggregation layer respectively to obtain a first network mapping result corresponding to each high-level query image block output by the high-level aggregation layer; determining the corresponding relationship between the high-level query image blocks and the first network mapping results as the first query mapping table; inputting each low-level query image block into a low-level aggregation layer respectively to obtain a second network output result corresponding to each low-level query image block output by the low-level aggregation layer; and determining the corresponding relationship between the low-level query image blocks and the second network mapping results as the second query mapping table.

[0147] Figure 11 The structural schematic diagram of an electronic device is exemplified as Figure 11 shown. The electronic device may include: a processor 1110, a communication interface 1120, a memory 1130, and a communication bus 1140. Among them, the processor 1110, the communication interface 1120, and the memory 1130 complete communication with each other through the communication bus 1140. The processor 1110 may call the logical instructions in the memory 1130 to execute the HDR image generation method provided by each of the above method embodiments.

[0148] In addition, when the logical instructions in the above-mentioned memory 1130 are implemented in the form of a software functional unit and sold or used as an independent product, they may be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, may be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The foregoing storage medium includes: various media such as a USB flash drive, a mobile hard disk, a read-only memory (ROM, Read-Only Memory), a random access memory (RAM, Random Access Memory), a magnetic disk, or an optical disc that can store program codes.

[0149] On the other hand, the present invention also provides a computer program product. The computer program product includes a computer program that can be stored on a computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the HDR image generation method provided by any one of the above method embodiments.

[0150] On the other hand, the present invention also provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, it is configured to execute the HDR image generation method provided in any of the above method embodiments.

[0151] Exemplarily, the computer-readable storage medium includes a non-transitory computer-readable storage medium.

[0152] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. A person of ordinary skill in the art can understand and implement it without creative efforts.

[0153] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solution, in essence, or the part that contributes to the prior art can be embodied in the form of a software product. The computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disc, etc., and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.

[0154] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them. Although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments or perform equivalent replacements for some of the technical features. However, these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A method for generating an HDR image, characterized in that, Including: Obtain the high - level image and the low - level image of the LDR image to be processed; Perform inverse tone mapping on the high - level image and the low - level image based on the inverse tone mapping network model to obtain the first mapped image corresponding to the high - level image and the second mapped image corresponding to the low - level image; wherein, the inverse tone mapping network model is obtained by training an initial inverse tone mapping network model based on the high - level sample image and the low - level sample image of the sample LDR image and the HDR label image corresponding to the sample LDR image; Merge the first mapped image and the second mapped image to generate the target HDR image corresponding to the LDR image to be processed.

2. The HDR image generation method according to claim 1, wherein The performing inverse tone mapping on the high - level image and the low - level image based on the inverse tone mapping network model to obtain the first mapped image corresponding to the high - level image and the second mapped image corresponding to the low - level image includes: Use the query mapping table corresponding to the inverse tone mapping network model to determine the first mapped image corresponding to the high - level image and the second mapped image corresponding to the low - level image; Wherein, the query mapping table is used to store the correspondence between each preset query image block and the network mapping result; the size of the preset query image block is determined based on the receptive field of the spatial enhancement layer in the inverse tone mapping network model.

3. The HDR image generation method according to claim 2, wherein The inverse tone mapping network model includes a high - level sub - network and a low - level sub - network; the query mapping table includes a first query mapping table corresponding to the high - level aggregation layer of the high - level sub - network and a second query mapping table corresponding to the low - level aggregation layer of the low - level sub - network; The using the query mapping table corresponding to the inverse tone mapping network model to determine the first mapped image corresponding to the high - level image and the second mapped image corresponding to the low - level image includes: Input the high - level image into the high - level spatial enhancement layer of the high - level sub - network to obtain the first feature map output by the high - level spatial enhancement layer; Use the first query mapping table to determine the first expansion result of the high - level aggregation layer after expanding the receptive field of the first feature map, and determine the first mapped image based on the first expansion result; Input the low - level image into the low - level spatial enhancement layer of the low - level sub - network to obtain the second feature map output by the low - level spatial enhancement layer; Use the second query mapping table to determine the second expansion result of the low - level aggregation layer after expanding the receptive field of the second feature map, and determine the second mapped image based on the second expansion result.

4. The HDR image generation method according to claim 3, wherein The using the first query mapping table to determine the first expansion result of the high - level aggregation layer after expanding the receptive field of the first feature map includes: Perform pixel - point dislocation fusion processing between at least two feature maps of the first feature map to obtain a third feature map; Use the first query mapping table to determine the network mapping result corresponding to each query image block in the third feature map to obtain the target output result of the high - level aggregation layer; Determine the first expansion result of the high - level aggregation layer after expanding the receptive field of the third feature map based on the target output result.

5. The HDR image generation method according to claim 4, wherein The high-level aggregation layer includes a first aggregation module. The first aggregation module includes at least one first horizontal aggregation layer and at least one first vertical aggregation layer, and the number of the first horizontal aggregation layers is the same as that of the first vertical aggregation layers. The process of performing pixel misalignment fusion processing between at least two feature maps on the first feature map to obtain a third feature map includes: Allocating the first feature map to each of the first horizontal aggregation layers and each of the first vertical aggregation layers; For each of the first horizontal aggregation layers, performing pixel misalignment fusion processing between at least two feature maps on the first feature map allocated to the first horizontal aggregation layer to obtain two third feature maps corresponding to the first horizontal aggregation layer; For each of the first vertical aggregation layers, performing pixel misalignment fusion processing between at least two feature maps on the first feature map allocated to the first vertical aggregation layer to obtain two third feature maps corresponding to the first vertical aggregation layer.

6. The HDR image generation method according to claim 5, wherein The high-level aggregation layer further includes at least one serially connected second aggregation module; the first query mapping table includes a first sub-query mapping table corresponding to the first aggregation module and second sub-query mapping tables respectively corresponding to the second aggregation modules; the process of using the first query mapping table to determine the network mapping result corresponding to each query image block in the third feature map to obtain the target output result of the high-level aggregation layer includes: For each of the first horizontal aggregation layers, using the first sub-query mapping table to determine the network mapping result corresponding to each query image block in the two third feature maps corresponding to the first horizontal aggregation layer to obtain a first sub-output result of the first horizontal aggregation layer; For each of the first vertical aggregation layers, using the first sub-query mapping table to determine the network mapping result corresponding to each query image block in the two third feature maps corresponding to the first vertical aggregation layer to obtain a second sub-output result of the first vertical aggregation layer; Jointly determining each of the first sub-output results and each of the second sub-output results as a first output result of the first aggregation module, and determining a fourth feature map based on the first output result and the first feature map; For each of the second aggregation modules, using the second sub-query mapping table corresponding to the second aggregation module to determine a second output result of the second aggregation module after expanding the receptive field of a fifth feature map; the fifth feature map is determined based on the output result of the previous aggregation module of the second aggregation module, and the input of the first second aggregation module is the fourth feature map; Determining the second output result of the last second aggregation module as the target output result of the high-level aggregation layer.

7. The HDR image generation method according to any one of claims 3 to 6, characterized in that, The process of determining the first mapping image based on the first expansion result includes: Fusing the first expansion result and the high-level image to obtain the first mapping image.

8. The HDR image generation method according to any one of claims 3 to 6, characterized in that The first query mapping table and the second query mapping table are established based on the following steps: Obtaining the preset query image block, where the preset query image block includes a high-level query image block and a low-level query image block; Input each of the high-level query image blocks into the high-level aggregation layer to obtain a first network mapping result corresponding to each of the high-level query image blocks output by the high-level aggregation layer; Determine the correspondence between the high-level query image blocks and the first network mapping results as the first query mapping table; Input each of the low-level query image blocks into the low-level aggregation layer to obtain a second network output result corresponding to each of the low-level query image blocks output by the low-level aggregation layer; Determine the correspondence between the low-level query image blocks and the second network mapping results as the second query mapping table.

9. An HDR image generation device, characterized in that, Comprising: An acquisition module, configured to acquire a high-level image and a low-level image of the LDR image to be processed; A mapping module, configured to perform inverse tone mapping on the high-level image and the low-level image based on an inverse tone mapping network model to obtain a first mapped image corresponding to the high-level image and a second mapped image corresponding to the low-level image; wherein, the inverse tone mapping network model is obtained by training an initial inverse tone mapping network model based on a high-level sample image and a low-level sample image of a sample LDR image and an HDR label image corresponding to the sample LDR image; A generation module, configured to merge the first mapped image and the second mapped image to generate a target HDR image corresponding to the LDR image to be processed.

10. An electronic device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, wherein, When the processor executes the computer program, the HDR image generation method according to any one of claims 1 to 8 is implemented.