Dynamic range compression method and device of fusion neural network

By converting image data into the logarithmic domain and using neural networks to enhance brightness and mid-to-high frequency information, combined with the DRC algorithm, the problem of insufficient image brightness and detail under low light or backlight conditions is solved, achieving higher quality image processing results.

CN120416489APending Publication Date: 2025-08-01SHANGHAI FULLHAN MICROELECTRONICS
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510764363.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-10
Publication Date
2025-08-01

AI Technical Summary

Technical Problem

Existing dynamic range compression algorithms suffer from poor image brightness and detail under low light or backlight conditions, making it difficult to meet the needs of security applications.

Method used

A fusion neural network approach is used to convert image data into the logarithmic domain and extract brightness guidance mapping information and mid-to-high frequency information using the neural network. Low-light enhancement and positive and negative enhancement processing are performed, and dynamic range compression is combined with the DRC algorithm.

Benefits of technology

It adaptively enhances the brightness and detail of images in low-light/backlit areas, improving image quality and making it suitable for object tracking, recognition, and detection tasks in security applications.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120416489A_ABST
    Figure CN120416489A_ABST
Patent Text Reader

Abstract

The invention provides a dynamic range compression method and device fused with a neural network. The method comprises the following steps: acquiring Y-format image data; converting the image data in the Y format into image data in a logarithm domain; converting the image data of the logarithm domain into brightness guidance mapping information and medium-high frequency information of the image by using a neural network; performing low-light enhancement on the image data of the logarithm domain by using the brightness guidance mapping information to obtain image data after low-light enhancement; performing positive and negative enhancement processing on the medium-high frequency information of the image to obtain enhanced medium-high frequency information; adding and fusing the low-light enhanced image data and the enhanced medium-high frequency information to obtain neural network enhanced image data; and compressing the dynamic range of the image data enhanced by the neural network by using a DRC algorithm to obtain primary output image data. According to the scheme, the brightness and detail performance of the image in the low-light / backlight area can be adaptively enhanced, and hardware implementation is facilitated.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image processing, and particularly relates to a dynamic range compression method and device integrating a neural network. Background Art

[0002] As Figure 1 shown, dynamic range compression (DRC) simulates the characteristics of the human eye visual system, compresses the high-dynamic range image obtained by the camera lens and sensor into a low-dynamic range image through the ISP (Image Signal Processor) system to adapt to traditional display devices, and expands the gray level range of the part that the human eye cares about, thereby achieving the purpose of improving the picture effect; the purpose of DRC is to enable the observer of the real scene and the observer of the display device to obtain similar visual experiences.

[0003] Currently, the mainstream traditional DRC algorithms can be divided into two types: the method based on histogram equalization and the method based on the Retinex theory. The latter has received relatively more attention. The Retinex theory believes that an image can be decomposed into two parts, as shown in formula (1):

[0004] I(x,y) = L(x,y) × R(x,y) (1)

[0005] where I is the image, L is the illumination component, R is the reflection component, and (x,y) is the coordinate position of the pixel point. The traditional DRC algorithm based on the Retinex theory generally realizes the dynamic range compression and contrast enhancement of the image by nonlinearly compressing the illumination component of the image and enhancing the reflection component of the image at the same time.

[0006] However, in practical applications, especially in the security field, the effects concerned by the market are sometimes not exactly the same as what the human eye actually perceives and what the camera / mobile phone captures. For example, due to inevitable environmental or technical limitations, images are often taken under suboptimal lighting conditions and are easily interfered by backlight / weak light, resulting in the degradation of the image information expression ability, such as the dimness of backlit human faces. The security application hopes that in such scenarios, images that are brighter and more detailed than the actual senses can also be presented, which is also beneficial for deploying higher-level tasks, such as object tracking, recognition, and detection. When the traditional DRC algorithm nonlinearly compresses the illumination component of the image, due to the relatively fixed non-linear mapping method, the brightness and detail performance of the image in the low-light / backlight area are poor. Summary of the Invention

[0007] The present invention provides a dynamic range compression method and device integrating a neural network to solve the technical problem that the brightness and details of an image are poorly presented in low-light / backlight areas.

[0008] To solve the above technical problem, the present invention provides a dynamic range compression method integrating a neural network, including the following steps:

[0009] S1. Obtain image data in Y format;

[0010] S2. Convert the image data in Y format into image data in the logarithmic domain;

[0011] S3. Use a neural network to convert the image data in the logarithmic domain into brightness guidance mapping information and mid-high frequency information of the image;

[0012] S4. Use the brightness guidance mapping information to perform low-light enhancement on the image data in the logarithmic domain to obtain low-light enhanced image data;

[0013] S5. Perform positive and negative enhancement processing on the mid-high frequency information of the image to obtain enhanced mid-high frequency information;

[0014] S6. Add and fuse the low-light enhanced image data and the enhanced mid-high frequency information to obtain image data enhanced by the neural network;

[0015] S7. Use the DRC algorithm to compress the dynamic range of the image data enhanced by the neural network to obtain primary output image data.

[0016] Preferably, step S3 specifically includes the following steps: successively process the image data in the logarithmic domain through a first convolutional layer and a leaky rectified linear unit activation function to obtain first-layer image data;

[0017] Successively process the first-layer image data through a second convolutional layer and a leaky rectified linear unit activation function to obtain second-layer image data;

[0018] Successively process the first-layer image data through a second convolutional layer and a leaky rectified linear unit activation function to obtain second-layer image data;

[0019] Successively process the second-layer image data through a third convolutional layer and a leaky rectified linear unit activation function to obtain third-layer image data;

[0020] Successively process the third-layer image data through a fourth convolutional layer and a leaky rectified linear unit activation function to obtain fourth-layer image data;

[0021] After splicing the fourth-layer image data and the third-layer image data, they are successively processed by a fifth convolutional layer and a leaky rectified linear unit activation function to obtain fifth-layer image data;

[0022] After splicing the fifth-layer image data and the second-layer image data, they are successively processed by a sixth convolutional layer and a leaky rectified linear unit activation function to obtain sixth-layer image data;

[0023] After splicing the sixth-layer image data and the first-layer image data, they are successively processed by a seventh convolutional layer and a hyperbolic tangent activation function to obtain the brightness guidance mapping information;

[0024] After splicing the fourth-layer image data and the third-layer image data, they are successively processed by an eighth convolutional layer and a rectified linear activation function to obtain eighth-layer image data;

[0025] The eighth-layer image data is successively processed by a ninth convolutional layer and a rectified linear activation function to obtain ninth-layer image data;

[0026] The ninth-layer image data is successively processed by a tenth convolutional layer and a rectified linear activation function to obtain the mid-high frequency information of the image.

[0027] Preferably, step S1 specifically includes the following steps: Determine whether the format of the input original image is the Y format. If it is, execute step S2; if not, convert the format of the original image to the Y format.

[0028] Preferably, step S2 specifically includes the following steps: Convert the image data in the Y format to the image data in the logarithmic domain through the formula LogY = Scale × log2(Y), where LogY represents the image data in the logarithmic domain, Scale represents the gain term, log2 represents the logarithmic operation with base 2, and Y represents the image data in the Y format.

[0029] Preferably, step S4 specifically includes the following steps: Normalize the brightness guidance mapping information and the image data in the logarithmic domain respectively; Through the formula LogY n = LogY n-1 + AiLumaMap n × LogY n-1 ×(1 - LogY n-1 ) Perform low-light enhancement on the normalized image data in the logarithmic domain, where LogY n represents the image data after low-light enhancement, that is, the image data in the logarithmic domain after n iterations, n represents the number of iterations, AiLumaMap n represents the normalized brightness guidance mapping information, LogY n-1Represents the image data in the logarithmic domain after n-1 iterations.

[0030] Preferably, the value range of n is [1, 8].

[0031] Preferably, step S5 specifically includes the following steps: Through the formula Perform positive and negative enhancement processing on the mid-high frequency information of the image to obtain the enhanced mid-high frequency information, where AiLogDetailE represents the enhanced mid-high frequency information, AiLogDetail represents the mid-high frequency information of the image, PosStr represents the positive enhancement intensity parameter, and NegStr represents the negative enhancement intensity parameter.

[0032] Preferably, after step S7, there is also step S8: Convert the format of the primary output image data to obtain the image data in the original format.

[0033] Preferably, step S8 specifically includes the following steps: Through the formula Convert the format of the primary output image data to obtain the image in the original format, where ImageOut represents the image data in the original format, ImageIn represents the original image data, Yo represents the primary output image data, and Y represents the image data in the Y format.

[0034] The present invention also provides a dynamic range compression device integrating a neural network, including the following units:

[0035] Brightness domain conversion unit, used to obtain image data in the Y format;

[0036] Logarithmic domain conversion unit, used to convert the image data in the Y format into image data in the logarithmic domain;

[0037] Neural network unit, used to convert the image data in the logarithmic domain into brightness guidance mapping information and the mid-high frequency information of the image by using a neural network;

[0038] Illumination component iteration unit, used to perform low-light enhancement on the image data in the logarithmic domain by using the brightness guidance mapping information to obtain the low-light enhanced image data;

[0039] Reflection component enhancement unit, used to perform positive and negative enhancement processing on the mid-high frequency information of the image to obtain the enhanced mid-high frequency information;

[0040] Fusion unit, used to add and fuse the low-light enhanced image data and the enhanced mid-high frequency information to obtain the neural network enhanced image data;

[0041] A DRC unit is used to compress the dynamic range of the image data enhanced by the neural network using the DRC algorithm to obtain primary output image data.

[0042] A method and device for dynamic range compression integrating a neural network provided by the present invention first convert Y-format image data into logarithmic-domain image data, then use the neural network to convert the logarithmic-domain image data into luminance guidance mapping information and mid-high frequency information of the image, then use the luminance guidance mapping information to perform low-light enhancement on the logarithmic-domain image data, and perform positive and negative enhancement processing on the mid-high frequency information of the image. Finally, use the DRC algorithm to compress the dynamic range of the image data enhanced by the neural network, so that the brightness and detail performance of the image in low-light / backlight areas can be enhanced adaptively, and it is beneficial to hardware implementation. BRIEF DESCRIPTION OF THE DRAWINGS

[0043] Figure 1 is a schematic structural diagram of an image dynamic range compression processing system in the prior art simulating the human visual system.

[0044] Figure 2 is a flowchart of a method for dynamic range compression integrating a neural network provided by an embodiment of the present invention.

[0045] Figure 3 is a schematic flowchart of a DRC unit performing dynamic range compression provided by an embodiment of the present invention.

[0046] Figure 4 is a schematic structural diagram of a device for dynamic range compression integrating a neural network provided by an embodiment of the present invention.

[0047] Figure 5 is a schematic structural diagram of a neural network unit provided by an embodiment of the present invention.

[0048] Figure 6 is a schematic structural diagram of a DRC unit provided by an embodiment of the present invention.

[0049] Figure 7 is a mapping relationship diagram between a base layer YBase restored to the luminance domain and a mapped base layer YBaseC provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0050] To make the objectives, advantages, and features of the present invention clearer, the following further describes in detail a method and device for dynamic range compression integrating a neural network proposed by the present invention with reference to the accompanying drawings. It should be noted that the accompanying drawings are all in very simplified forms and use non-precise scales, only for the purpose of facilitating and clearly assisting in explaining the objectives of the embodiments of the present invention.

[0051] In the description of the present invention, the qualifiers such as "first", "second", etc. are added for convenience of description and reference, and should not be construed as indicating or implying relative importance or implicitly specifying the number of the indicated technical features. Thus, the features defined with qualifiers such as "first", "second", etc. may explicitly or implicitly include one or more of such features.

[0052] As Figure 2 shown, this embodiment provides a dynamic range compression method integrating a neural network, including the following steps:

[0053] S1. Obtain image data in Y format. An image in Y format refers to a black-and-white grayscale image. This step specifically includes the following steps: Determine whether the format of the input original image is Y format. If it is, execute step S2; if not, convert the format of the original image to Y format. If the input image data is in Raw format, it needs to be first converted to RGB format and then to Y format. Here, a complex Demosaic algorithm can be used for conversion, or a simple interpolation algorithm can be used for conversion; if the input image data is in RGB format, it is directly converted to image data in Y format, and its calculation formulas (2 - 3) are as follows:

[0054] Y′(x,y) = W r ×R(x,y) + W g ×G(x,y) + W b ×B(x,y)

[0055] And W r +W g +W b = 1 (2)

[0056] Y(x,y) = (1 - W y )×Y ′ (x,y) + W y ×max(R(x,y), G(x,y), B(x,y)) (3)

[0057] where Y′ is the initial image data in Y format, Y is the image data in Y format, R, G, and B are the RGB components of the input image data, i.e., the original image data; W r , W g , W b are the weighted weight parameters of the three components (the sum of which must be 1); max is a logical operation function representing taking the maximum value; W y is the maximum value among the three weighted weight parameters, and (x,y) is the coordinate of the pixel.

[0058] S2. Convert the image data in Y format into image data in the logarithmic domain. The reasons for converting the image data in Y format into image data in the logarithmic domain are mainly as follows: First, because the human visual system's perception of brightness changes is based on relative contrast changes, and its changes are non-linear. Performing Log domain conversion, that is, logarithmic domain conversion, can effectively expand the dynamic range of the dark area while compressing the dynamic range of the bright area, thus better simulating the non-linear photosensitive system of the human eye. Second, performing addition operations in the Log domain can be equivalent to multiplication operations in the normal brightness domain, which can reduce the calculation logic and save hardware overhead. The calculation formula (4) for Log domain conversion is as follows:

[0059] LogY = Scale × log2(Y) (4)

[0060] Among them, LogY represents the image data in the logarithmic domain, Scale represents the gain term (used to adjust the fixed-point bit width of the image data), log2 represents the logarithmic operation with base 2, and Y represents the image data in Y format.

[0061] S3. Use a neural network to convert the image data in the logarithmic domain into brightness guidance mapping information and the mid-high frequency information of the image. A neural network usually contains multiple convolutional layers (Convolution, abbreviated as Conv). The calculation formula (5) of the convolutional layer is as follows:

[0062]

[0063] Among them, A d,i,j is the output feature map of the current convolutional layer; W d,m,n is the convolutional kernel weight of the current convolutional layer; X d,i+m,j+n is the input feature map of the current convolutional layer, and i, j are its pixel indices; Bias is the bias term of the current convolutional layer; D represents the depth of the input feature map of the current convolutional layer; d, m, n are the channel indices on the three dimensions (length, width, and height) of the input feature map; F represents the size of the convolutional kernel of the current convolutional layer (F of each convolutional layer usually takes 3). Each convolutional layer is connected in series with an activation function. The three activation functions are as follows:

[0064]

[0065]

[0066] Among them, ReLU is the rectified linear unit activation function, Tanh is the hyperbolic tangent activation function, Leaky ReLU is the leaky rectified linear unit activation function, X is the input value of the activation function, that is, the output feature map of the current convolutional layer, and α is the negative slope parameter term of the current LeakyReLU activation layer (α of each Leaky ReLU activation layer can all be 0.02).

[0067] Preferably, as Figure 5 shown, step S3 specifically includes the following steps: successively passing the image data in the logarithmic domain through a first convolutional layer and a leaky rectified linear unit activation function Leaky ReLU to obtain first-layer image data;

[0068] successively passing the first-layer image data through a second convolutional layer and a leaky rectified linear unit activation function LeakyReLU to obtain second-layer image data;

[0069] successively passing the first-layer image data through a second convolutional layer and a leaky rectified linear unit activation function LeakyReLU to obtain second-layer image data;

[0070] successively passing the second-layer image data through a third convolutional layer and a leaky rectified linear unit activation function Leaky ReLU to obtain third-layer image data;

[0071] successively passing the third-layer image data through a fourth convolutional layer and a leaky rectified linear unit activation function Leaky ReLU to obtain fourth-layer image data;

[0072] successively passing the concatenated fourth-layer image data and third-layer image data through a fifth convolutional layer and a leaky rectified linear unit activation function Leaky ReLU to obtain fifth-layer image data; wherein, concatenation means aligning the length and width directions of two image layers and stacking them in the height direction.

[0073] successively passing the concatenated fifth-layer image data and second-layer image data through a sixth convolutional layer and a leaky rectified linear unit activation function Leaky ReLU to obtain sixth-layer image data;

[0074] successively passing the concatenated sixth-layer image data and first-layer image data through a seventh convolutional layer and a hyperbolic tangent activation function Tanh to obtain the luminance guidance mapping information;

[0075] successively passing the concatenated fourth-layer image data and third-layer image data through an eighth convolutional layer and a rectified linear unit activation function ReLU to obtain eighth-layer image data;

[0076] successively passing the eighth-layer image data through a ninth convolutional layer and a rectified linear unit activation function ReLU to obtain ninth-layer image data;

[0077] successively passing the ninth-layer image data through a tenth convolutional layer and a rectified linear unit activation function ReLU to obtain the mid-high frequency information of the image.

[0078] S4. Use the luminance guidance mapping information to perform low-light enhancement on the image data in the logarithmic domain to obtain the low-light enhanced image data. This step is to perform low-light enhancement on the image data LogY in the logarithmic domain in step S2, that is, to increase the luminance of the image data LogY in the logarithmic domain. Specifically, it may include the following steps: Normalize the luminance guidance mapping information and the image data in the logarithmic domain respectively; Perform low-light enhancement on the normalized image data in the logarithmic domain through formula (9).

[0079] LogY n = LogY n-1 + AiLumaMap n × LogY n-1 × (1 - LogY n-1 ) (9)

[0080] Where LogY n represents the image data after low-light enhancement, that is, the image data in the logarithmic domain after n iterations. n represents the number of iterations. AiLumaMap n represents the normalized luminance guidance mapping information. LogY n-1 represents the image data in the logarithmic domain after n - 1 iterations. The physical meaning of n is to control the curvature of the non-linearity. Effectively, the larger n is, the more iterations are performed, and the stronger the enhancement effect. The value range of n is preferably [1, 8].

[0081] S5. Perform positive and negative enhancement processing on the mid-high frequency information of the image to obtain the enhanced mid-high frequency information. The process of this step is shown in formula (10):

[0082]

[0083] Where AiLogDetailE represents the enhanced mid-high frequency information, AiLogDetail represents the mid-high frequency information of the image, PosStr represents the positive enhancement intensity parameter, and NegStr represents the negative enhancement intensity parameter.

[0084] S6. Add and fuse the low-light enhanced image data and the enhanced mid-high frequency information to obtain the image data AiLogY enhanced by the neural network. The process of this step is shown in formula (11):

[0085] AiLogY = LogY n + AiLogDetailE (11)

[0086] S7. Compress the dynamic range of the image data enhanced by the neural network using the DRC algorithm to obtain primary output image data. The DRC algorithm is an existing traditional DRC algorithm that compresses the dynamic range of image data based on the Retinex theory. For ease of description, as Figure 3 and Figure 6 shown, this step is described through the following several sub-units: a hierarchical filtering unit 171, a luminance domain restoration unit 172, a non-linear compression unit 173, a detail enhancement unit 174, and an image format restoration unit 175.

[0087] The hierarchical filtering unit 171 is used to execute step S71: decompose the image data AiLogY enhanced by the neural network into a light component (hereinafter referred to as the base layer, denoted as LogYBase) and a reflection component (hereinafter referred to as the detail layer, LogYDetail). The specific implementation method of this unit is to perform edge-preserving filtering processing with a large perception domain on the image data input to this unit to extract the main luminance information of the image. The methods are not limited to: guided filtering, bilateral filtering and other algorithms. In the present invention, bilateral filtering (Bilateral Filter, abbreviated as BF) is taken as an example, and its calculation process is shown in formulas (12) and (13):

[0088]

[0089] Among them, I represents the pixel value; p represents the spatial position of the current point; q represents the spatial position of the neighborhood point; W p represents the neighborhood pixel weight centered on point p; S represents the neighborhood space; and both represent the Gaussian function (GaussianFunction), σ s and σ r respectively correspond to the standard deviation control coefficients of the spatial domain (SpatialDomain) and the range domain (Range Domain); |||| represents the L2 norm (Euclidean distance). The Gaussian function is shown in formulas (14) and (15):

[0090]

[0091] Substitute AiLogY into I in formula (12), and the image base layer information LogYBase after BF processing can be obtained; calculate the residual between it and the input AiLogY, and the image detail layer information LogYDetail can be obtained. The calculation process is shown in formula (16):

[0092] LogYDetail = AilogY - LogYBase (16)

[0093] The luminance domain restoration unit 172 is used to perform step S72: restoring the image data from the Log domain back to the luminance domain, which is essentially the inverse function of the logarithmic function, that is, the exponential function. Its calculation formula (17) is as follows:

[0094]

[0095] where YBase represents the base layer restored to the luminance domain.

[0096] The non-linear compression unit 173 is used to perform step S73: performing non-linear mapping on the base layer YBase restored to the luminance domain to obtain the mapped base layer YBaseC, so as to expand the dynamic range of the dark area and compress the dynamic range of the bright area at the same time. The mapping relationship is as Figure 7 shown. The specific implementation method can be a Look-Up-Table (LUT for short), and its calculation formula (18) is as follows:

[0097] YBaseC = LUT(YBase) (18)

[0098] The detail enhancement unit 174 is used to perform step S74: performing positive and negative enhancement processing on the reflection component, that is, the detail layer LogYDetail. Its calculation process is as shown in formula (19):

[0099]

[0100] where LogYDetailE is the enhanced detail layer, PosStrC is the positive enhancement intensity parameter; NegStrC is the negative enhancement intensity parameter.

[0101] After the enhanced detail layer LogYDetailE is processed by the luminance domain restoration unit 172, it is restored from the Log domain back to the luminance domain to obtain the detail layer YDetail restored to the luminance domain.

[0102] The mapped base layer YBaseC and the detail layer YDetail restored to the luminance domain are added together to obtain the primary output image data Yo. The calculation process is as shown in formula (20):

[0103] Yo = YBaseC + YDetail (20)

[0104] The image format restoration unit 175 is used to perform step S75, that is, S8: converting the format Yo of the primary output image data to obtain the image data ImageOut in the original format (that is, Figure 4 the original format of the input image data). Specifically, the primary output image data Yo can be restored to the original Raw format or RGB format. Its calculation process is as shown in formula (21):

[0105]

[0106] Among them, ImageOut represents the image data in the original format, ImageIn represents the original image data, Yo represents the primary output image data, and Y represents the image data in the Y format.

[0107] A dynamic range compression method of a fusion neural network provided in this embodiment first converts the image data in the Y format into image data in the logarithmic domain, then uses the neural network to convert the image data in the logarithmic domain into luminance guidance mapping information and the mid-high frequency information of the image, then uses the luminance guidance mapping information to perform low-light enhancement on the image data in the logarithmic domain, and performs positive and negative enhancement processing on the mid-high frequency information of the image. Finally, the DRC algorithm is used to compress the dynamic range of the image data enhanced by the neural network, so that the brightness and detail performance of the image in the low-light / backlight area can be adaptively enhanced, and it is beneficial to hardware implementation.

[0108] As Figure 4 shown, based on the same technical concept as the above-mentioned dynamic range compression method of a fusion neural network, this embodiment provides a dynamic range compression device of a fusion neural network, including the following units:

[0109] The luminance domain conversion unit 11 is used to obtain the image data in the Y format;

[0110] The logarithmic domain conversion unit 12 is used to convert the image data in the Y format into image data in the logarithmic domain;

[0111] The neural network unit 13 is used to use the neural network to convert the image data in the logarithmic domain into luminance guidance mapping information and the mid-high frequency information of the image;

[0112] The illumination component iteration unit 14 is used to perform low-light enhancement on the image data in the logarithmic domain by using the luminance guidance mapping information to obtain the image data after low-light enhancement;

[0113] The reflection component enhancement unit 15 is used to perform positive and negative enhancement processing on the mid-high frequency information of the image to obtain the enhanced mid-high frequency information;

[0114] The fusion unit 16 is used to add and fuse the image data after low-light enhancement and the enhanced mid-high frequency information to obtain the image data enhanced by the neural network;

[0115] The DRC unit 17 is used to compress the dynamic range of the image data enhanced by the neural network by using the DRC algorithm to obtain the primary output image data.

[0116] A dynamic range compression device integrating a neural network first converts Y-format image data into logarithmic-domain image data, then uses the neural network to convert the logarithmic-domain image data into luminance guidance mapping information and the mid-high frequency information of the image, then uses the luminance guidance mapping information to perform low-light enhancement on the logarithmic-domain image data, and performs positive and negative enhancement processing on the mid-high frequency information of the image. Finally, the DRC algorithm is used to compress the dynamic range of the image data enhanced by the neural network. In this way, the brightness and detail performance of the image in low-light / backlight areas can be enhanced adaptively, and it is beneficial for hardware implementation.

[0117] In summary, a dynamic range compression method and device integrating a neural network provided by the present invention first convert Y-format image data into logarithmic-domain image data, then use the neural network to convert the logarithmic-domain image data into luminance guidance mapping information and the mid-high frequency information of the image, then use the luminance guidance mapping information to perform low-light enhancement on the logarithmic-domain image data, and perform positive and negative enhancement processing on the mid-high frequency information of the image. Finally, the DRC algorithm is used to compress the dynamic range of the image data enhanced by the neural network. In this way, the brightness and detail performance of the image in low-light / backlight areas can be enhanced adaptively, and it is beneficial for hardware implementation.

[0118] The above description is only a description of the preferred embodiments of the present invention, and does not limit the scope of the present invention in any way. Any changes and modifications made by those of ordinary skill in the art according to the above disclosure shall fall within the protection scope of the present invention.

Claims

1. A dynamic range compression method integrating a neural network, characterized in that It includes the following steps: S1. Obtain image data in Y format; S2. Convert the image data in Y format into image data in the logarithmic domain; S3. Use a neural network to convert the image data in the logarithmic domain into luminance guidance mapping information and mid-high frequency information of the image; S4. Use the luminance guidance mapping information to perform low-light enhancement on the image data in the logarithmic domain to obtain low-light enhanced image data; S5. Perform positive and negative enhancement processing on the mid-high frequency information of the image to obtain enhanced mid-high frequency information; S6. Add and fuse the low-light enhanced image data and the enhanced mid-high frequency information to obtain image data enhanced by the neural network; S7. Use the DRC algorithm to compress the dynamic range of the image data enhanced by the neural network to obtain primary output image data.

2. The dynamic range compression method integrating a neural network according to claim 1, characterized in that Step S3 specifically includes the following steps: successively process the image data in the logarithmic domain through a first convolutional layer and a leaky rectified linear unit activation function to obtain first-layer image data; successively process the first-layer image data through a second convolutional layer and a leaky rectified linear unit activation function to obtain second-layer image data; successively process the first-layer image data through a second convolutional layer and a leaky rectified linear unit activation function to obtain second-layer image data; successively process the second-layer image data through a third convolutional layer and a leaky rectified linear unit activation function to obtain third-layer image data; successively process the third-layer image data through a fourth convolutional layer and a leaky rectified linear unit activation function to obtain fourth-layer image data; concatenate the fourth-layer image data and the third-layer image data and then successively process them through a fifth convolutional layer and a leaky rectified linear unit activation function to obtain fifth-layer image data; concatenate the fifth-layer image data and the second-layer image data and then successively process them through a sixth convolutional layer and a leaky rectified linear unit activation function to obtain sixth-layer image data; concatenate the sixth-layer image data and the first-layer image data and then successively process them through a seventh convolutional layer and a hyperbolic tangent activation function to obtain the luminance guidance mapping information; concatenate the fourth-layer image data and the third-layer image data and then successively process them through an eighth convolutional layer and a rectified linear activation function to obtain eighth-layer image data; successively process the eighth-layer image data through a ninth convolutional layer and a rectified linear activation function to obtain ninth-layer image data; successively process the ninth-layer image data through a tenth convolutional layer and a rectified linear activation function to obtain the mid-high frequency information of the image.

3. The dynamic range compression method integrating a neural network according to claim 1, characterized in that Step S1 specifically includes the following steps: determine whether the format of the input original image is in Y format. If so, execute step S2; if not, convert the format of the original image into Y format.

4. A dynamic range compression method integrating a neural network, characterized in that, Step S2 specifically includes the following steps: convert the image data in Y format into image data in the logarithmic domain through the formula LogY = Scale × log2(Y), where LogY represents the image data in the logarithmic domain, Scale represents the gain term, log2 represents the logarithmic operation with base 2, and Y represents the image data in Y format.

5. The dynamic range compression method integrating a neural network according to claim 1, characterized in that Step S4 specifically includes the following steps: perform normalization processing on the luminance guidance mapping information and the image data in the logarithmic domain respectively; through the formula LogY n = LogY n-1 + AiLumaMap n × LogY n-1 × (1 - LogY n-1 ) perform low-light enhancement on the normalized image data in the logarithmic domain, where LogY n represents the image data after low-light enhancement, that is, the image data in the logarithmic domain after n iterations, n represents the number of iterations, AiLumaMap n represents the normalized luminance guidance mapping information, and LogY n-1 represents the image data in the logarithmic domain after n - 1 iterations.

6. The dynamic range compression method integrating a neural network according to claim 5, characterized in that, The value range of n is [1, 8].

7. The dynamic range compression method integrating a neural network according to claim 1, wherein Step S5 specifically includes the following steps: Through the formula perform positive and negative enhancement processing on the mid-high frequency information of the image to obtain the enhanced mid-high frequency information, where AiLogDetailE represents the enhanced mid-high frequency information, AiLogDetail represents the mid-high frequency information of the image, PosStr represents the positive enhancement intensity parameter, and NegStr represents the negative enhancement intensity parameter.

8. A dynamic range compression method integrating a neural network, characterized in that, After step S7, there is also step S8: converting the format of the primary output image data to obtain the image data in the original format.

9. The dynamic range compression method of a fusion neural network according to claim 8, wherein Step S8 specifically includes the following steps: Through the formula convert the format of the primary output image data to obtain an image in the original format, where ImageOut represents the image data in the original format, ImageIn represents the original image data, Yo represents the primary output image data, and Y represents the image data in the Y format.

10. A dynamic range compression device integrating a neural network, characterized in that, It includes the following units: A luminance domain conversion unit for obtaining image data in Y format; A logarithmic domain conversion unit for converting the image data in Y format into image data in the logarithmic domain; A neural network unit for converting the image data in the logarithmic domain into luminance guidance mapping information and mid-high frequency information of the image by using a neural network; A lighting component iteration unit for performing low-light enhancement on the image data in the logarithmic domain by using the luminance guidance mapping information to obtain low-light enhanced image data; A reflection component enhancement unit for performing positive and negative enhancement processing on the mid-high frequency information of the image to obtain enhanced mid-high frequency information; A fusion unit for adding and fusing the low-light enhanced image data and the enhanced mid-high frequency information to obtain neural network enhanced image data; A DRC unit for compressing the dynamic range of the neural network enhanced image data by using the DRC algorithm to obtain primary output image data.