Image Enhancement Method and Device

By reconstructing and fusing the basic information and detailed information of images under low-light conditions, the problem of difficulty in achieving image brightness and contrast enhancement and denoising in the prior art is solved, and high-quality image enhancement effect is achieved.

CN115769247BActive Publication Date: 2025-06-10HUAWEI TECH CO LTD +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202080101508.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2020-07-27
Publication Date
2025-06-10
Estimated Expiration
2040-07-27

AI Technical Summary

Technical Problem

The prior art is difficult to simultaneously realize brightness and contrast enhancement of images and image denoising under low light conditions.

Method used

The low-frequency feature map of the to-process image is obtained through the first neural network and the basic information is reconstructed; the high-frequency feature map of the to-process image is obtained through the second neural network and the detailed information is reconstructed; and the basic information and the detailed information are fused to obtain the enhanced image.

Benefits of technology

It realizes the effective filtering of image noise while enhancing image brightness and contrast, and improves image quality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115769247B_ABST
    Figure CN115769247B_ABST
Patent Text Reader

Abstract

An image enhancement method and apparatus, relating to the field of image processing technology. This method can simultaneously enhance the brightness and / or contrast of an image and denoise the image. The method includes: obtaining a low-frequency feature map of the image to be processed through a first neural network, and determining a first image based on the low-frequency feature map; wherein, the first image includes the basic information of the reconstructed image to be processed, and the basic information includes the contour information of the image to be processed; then, obtaining a high-frequency feature map of the image to be processed through a second neural network, and determining a second image based on the high-frequency feature map; wherein, the second image includes the detailed information of the reconstructed image to be processed, and the detailed information includes at least one of the edges or textures of the image to be processed; and then, fusing the first image and the second image to obtain the enhanced image of the image to be processed.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of image processing technology, and in particular to an image enhancement method and device. Background Art

[0002] Low light conditions are very common in daily life, such as nighttime environments, sunset environments, etc. Therefore, images captured under low light conditions (such as surveillance images captured during public environment surveillance at night, or personal images captured at sunset, etc.) usually have a low signal-to-noise ratio and severe noise.

[0003] Generally, an image captured under low-light conditions can be adjusted to an image acceptable to the human eye by an image enhancement method, that is, the brightness and / or contrast of the image can be enhanced. However, when processing images captured under low-light conditions by traditional image enhancement methods, it is impossible to simultaneously achieve the effects of enhancing the brightness and / or contrast of the image and denoising the image. Therefore, how to simultaneously achieve the enhancement of the brightness and / or contrast of the image captured under low-light conditions and denoising of the image is a technical problem that needs to be solved urgently. Summary of the invention

[0004] The present application provides an image enhancement method and device, through which image brightness and / or contrast enhancement and image denoising can be achieved simultaneously.

[0005] To achieve the above objectives, this application provides the following technical solutions:

[0006] In a first aspect, the present application provides an image enhancement method, which is applied to an image enhancement device. The method includes: obtaining a low-frequency feature map of an image to be processed through a first neural network, and determining a first image based on the low-frequency feature map. The first image includes basic information of the reconstructed image to be processed, and the basic information includes contour information of the image to be processed. Then, a high-frequency feature map of the image to be processed is obtained through a second neural network, and a second image is determined based on the high-frequency feature map. The second image includes detailed information of the reconstructed image to be processed, and the detailed information includes at least one of an edge or texture of the image to be processed. Then, the first image and the second image are fused to obtain an enhanced image of the image to be processed.

[0007] In this way, the image enhancement method provided by the present application can denoise and reconstruct the base image by obtaining the low-frequency feature map of the image to be processed (wherein, denoising is because most of the noise information is high-frequency information, and reconstructing the base image according to the low-frequency features is equivalent to filtering out most of the high-frequency information, that is, denoising is achieved), and reconstruct the detail image by obtaining the high-frequency feature map of the image to be processed, and then fuse the base image and the detail image to obtain the enhanced image of the image to be processed. Through this method, while enhancing the brightness and / or contrast of the image, the noise in the image to be processed can be effectively filtered out.

[0008] In a possible design, the above-mentioned "obtaining the low-frequency feature map of the image to be processed through the first neural network" includes: using the first neural network to obtain the first feature map of the image to be processed, and then, using the first neural network to multiply the pixel value of the first pixel in the first feature map by the pixel value of the second pixel corresponding to the first pixel in the image to be processed to obtain the low-frequency feature map of the image to be processed. The above-mentioned "obtaining the high-frequency feature map of the image to be processed through the second neural network" includes: using the second neural network to invert the pixel value of each pixel in the aforementioned first feature map to obtain the second feature map. Then, using the second neural network to multiply the pixel value of the third pixel in the second feature map by the pixel value of the fourth pixel corresponding to the third pixel in the aforementioned first image, or using the second neural network to multiply the pixel value of the third pixel by the pixel value of the fifth pixel corresponding to the third pixel in the image to be processed to obtain the high-frequency feature map of the image to be processed. Among them, the network structures of the second neural network and the first neural network are the same.

[0009] Through this possible design, when the method of the present application obtains the high-frequency feature map of the image to be processed, it shares the parameters for obtaining the low-frequency feature map of the image to be processed, that is, the image enhancement device inverts the first feature map used to determine the image to be processed to obtain the second feature map used to determine the high-frequency feature map of the image to be processed. Through this design, a large number of convolution operations (that is, the convolution operations for obtaining the second feature map according to the image to be processed or the first image) can be reduced during the process of enhancing the image to be processed by the method of the present application, thereby saving computing power.

[0010] In another possible design, the above-mentioned "inversion" means: subtracting 1 from the pixel value of each pixel in the first feature map.

[0011] Through this possible design, the second feature map used to determine the high-frequency feature map of the image to be processed can be simply and quickly obtained based on the first feature map used to determine the image to be processed, thereby reducing a large number of convolution operations (that is, the convolution operations for obtaining the second feature map according to the image to be processed or the first image) and saving computing power.

[0012] In another possible design, the above-mentioned "determining the first image according to the low-frequency feature map" includes: according to the low-frequency feature map, using a third neural network to reconstruct the basic information of the image to be processed to obtain a third image. Then, enhancing the color and / or contrast of the third image by a constant α to obtain a fourth image. Then, multiplying the pixel value of the sixth pixel in the third image by the pixel value of the seventh pixel corresponding to the sixth pixel in the fourth image to obtain the first image.

[0013] In another possible design, the above-mentioned "constant α" is a preset constant, or the above-mentioned "constant α" is obtained through a third neural network.

[0014] In another possible design, the above-mentioned "determining the second image according to the high-frequency feature map" includes: according to the high-frequency feature map, using a fourth neural network to reconstruct the detailed information of the image to be processed to obtain a second image. Among them, the network structure of the fourth neural network is the same as that of the above-mentioned third neural network. The feature map used to obtain the high-frequency feature map in the fourth neural network is obtained by taking the inverse of the pixel value of each pixel in the feature map used to obtain the low-frequency feature map in the above-mentioned third neural network.

[0015] Through the above several possible designs, the fourth reconstruction network can share the parameters of the third reconstruction network, that is, by taking the inverse of the pixel value in the feature map used to obtain the low-frequency feature map in the third neural network, the feature map used to obtain the high-frequency feature map in the fourth neural network can be obtained. In this way, a large number of convolution operations (that is, the convolution operations for obtaining the feature map used to obtain the high-frequency feature map in the fourth neural network) can be reduced, thereby saving computing power.

[0016] In another possible design, the above-mentioned "fusing the first image and the second image to obtain the enhanced image of the image to be processed" includes: adding the pixel value of the eighth pixel in the first image to the pixel value of the ninth pixel corresponding to the eighth pixel in the second image to obtain the enhanced image of the image to be processed.

[0017] In a second aspect, the present application provides an image enhancement device.

[0018] In a possible design, the image enhancement device is used to execute any of the methods provided in the above first aspect. This application can divide the functional modules of the image enhancement device according to any of the methods provided in the above first aspect. For example, each functional module can be divided corresponding to each function, or two or more functions can be integrated into one processing module. Exemplarily, this application can divide the image enhancement device into an acquisition unit, a determination unit, a fusion unit, etc. according to functions. The descriptions of the possible technical solutions and beneficial effects executed by each of the above-divided functional modules can refer to the technical solutions provided in the above first aspect or its corresponding possible design, which will not be elaborated here.

[0019] In another possible design, the image enhancement device includes: a memory and one or more processors, and the memory and the processor are coupled. The memory is used to store computer instructions, and the processor is used to call the computer instructions to execute any of the methods provided in the first aspect and any of its possible design manners.

[0020] In a third aspect, this application provides a computer-readable storage medium, such as a non-transitory computer-readable storage medium. Computer programs (or instructions) are stored thereon. When the computer programs (or instructions) run on the image enhancement device, the image enhancement device is enabled to execute any of the methods provided in any of the possible implementation manners in the above first aspect.

[0021] In a fourth aspect, this application provides a computer program product, which when running on the image enhancement device, enables any of the methods provided in any of the possible implementation manners in the first aspect to be executed.

[0022] In a fifth aspect, this application provides a chip system, including: a processor, and the processor is used to call and run the computer program stored in the memory from the memory and execute any of the methods provided in the implementation manner in the first aspect.

[0023] It can be understood that any of the above-provided devices, computer storage media, computer program products, or chip systems, etc. can be applied to the corresponding methods provided above. Therefore, the beneficial effects that can be achieved can refer to the beneficial effects in the corresponding methods, which will not be elaborated here.

[0024] In this application, the name of the above image enhancement device does not constitute a limitation on the device or functional module itself. In actual implementation, these devices or functional modules can appear under other names. As long as the functions of each device or functional module are similar to those of this application and belong to the scope of the claims of this application and their equivalent technologies.

[0025] These aspects or other aspects of this application will be more clearly understood in the following description. Description of the Drawings

[0026] Figure 1 Schematic diagram of the hardware structure of an image enhancement device provided by an embodiment of the present application Figure 1 ;

[0027] Figure 2 Schematic diagram of the hardware structure of an image enhancement device provided by an embodiment of the present application Figure 2 ;

[0028] Figure 3 Schematic flowchart of the image enhancement method provided by an embodiment of the present application;

[0029] Figure 4 Network structure diagram of the ACE network provided by an embodiment of the present application;

[0030] Figure 5 Schematic diagram of corresponding pixels in different feature maps provided by an embodiment of the present application;

[0031] Figure 6 Schematic structure diagram of the first reconstruction network provided by an embodiment of the present application;

[0032] Figure 7 Schematic structure diagram of the first CDT module provided by an embodiment of the present application;

[0033] Figure 8 Schematic structure diagram of an image enhancement model provided by an embodiment of the present application;

[0034] Figure 9 Schematic flowchart of a method for training an image enhancement model provided by an embodiment of the present application;

[0035] Figure 10 Schematic structure diagram of an image enhancement device provided by an embodiment of the present application;

[0036] Figure 11 Schematic structure diagram of a chip system provided by an embodiment of the present application;

[0037] Figure 12 Schematic structure diagram of a computer program product provided by an embodiment of the present application. Detailed Embodiments

[0038] In the embodiments of the present application, words such as "exemplary" or "for example" are used to represent examples, illustrations or explanations. Any embodiment or design solution described as "exemplary" or "for example" in the embodiments of the present application should not be construed as being more preferred or having more advantages than other embodiments or design solutions. Rather, the use of words such as "exemplary" or "for example" is intended to present related concepts in a specific manner.

[0039] In the embodiments of the present application, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, features defined with "first" and "second" may explicitly or implicitly include one or more of such features. In the description of the present application, unless otherwise specified, the meaning of "a plurality" is two or more.

[0040] In the present application, the meaning of the term "at least one" is one or more, and the meaning of the term "a plurality" is two or more. For example, a plurality of second messages means two or more second messages. In this document, the terms "system" and "network" are often used interchangeably.

[0041] It should be understood that the terms used in the description of various examples herein are only for the purpose of describing specific examples and are not intended to be limiting. As used in the description of various examples and the appended claims, the singular forms "a", "an", and "the" are also intended to include the plural forms unless the context clearly indicates otherwise.

[0042] It should also be understood that the term "and / or" as used herein refers to and encompasses any and all possible combinations of one or more of the associated listed items. The term "and / or" describes the associative relationship of associated objects and indicates that three relationships may exist. For example, A and / or B may represent: A exists alone, A and B exist simultaneously, and B exists alone. In addition, the character " / " in the present application generally represents an "or" relationship between the associated objects before and after.

[0043] It should also be understood that in the various embodiments of the present application, the magnitude of the serial numbers of the various processes does not mean the order of execution, and the order of execution of each process should be determined by its function and internal logic and should not constitute any limitation to the implementation process of the embodiments of the present application.

[0044] It should be understood that determining B based on A does not mean determining B only based on A, and B can also be determined based on A and / or other information.

[0045] It should also be understood that the term "comprise" (also known as "includes", "including", "comprises", and / or "comprising") when used in this specification specifies the presence of the stated features, integers, steps, operations, elements, and / or components, but does not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or their groups.

[0046] It should also be understood that the term "if" can be interpreted to mean "when" or "upon" or "in response to determining" or "in response to detecting". Similarly, depending on the context, the phrase "if it is determined that..." or "if [the stated condition or event] is detected" can be interpreted to mean "when it is determined that..." or "in response to determining..." or "when [the stated condition or event] is detected" or "in response to detecting [the stated condition or event]".

[0047] It should be understood that the "one embodiment", "an embodiment", "a possible implementation" mentioned throughout the specification means that the specific features, structures or characteristics related to the embodiment or implementation are included in at least one embodiment of the present application. Therefore, the "in one embodiment" or "in an embodiment", "a possible implementation" that appear throughout the specification do not necessarily refer to the same embodiment. In addition, these specific features, structures or characteristics can be combined in one or more embodiments in any suitable manner.

[0048] The embodiments of the present application provide an image enhancement method. This method reconstructs the basic information of the low-frequency feature image of the image to be processed, reconstructs the details of the high-frequency feature map of the image to be processed, and then fuses the reconstructed images to obtain the enhanced image to be processed. The enhanced image obtained by this method can not only enhance the brightness and / or contrast of the image to be processed, but also filter out the noise of the image.

[0049] Among them, the image to be processed can be an image taken under low light conditions, or any image that needs to be enhanced. The embodiments of the present application do not limit this.

[0050] Among them, the above-mentioned basic information includes the contour information of the image to be processed, etc. The above-mentioned detail information includes at least one of the edges or textures of the image to be processed.

[0051] The embodiments of the present application provide an image enhancement device for performing the above-mentioned image enhancement method. This image enhancement device can be a terminal. Alternatively, this image enhancement device can be a server.

[0052] Among them, the above-mentioned terminal can be a portable device such as a mobile phone, a tablet computer, a wearable electronic device, etc., or a computing device such as a personal computer (PC), a personal digital assistant (PDA), a netbook, etc., or any other terminal device that can implement the embodiments of the present application. The present application does not limit this.

[0053] Among them, when the above image enhancement device is a terminal, the above image enhancement method can be implemented through an application installed on the terminal, such as a client application for processing images and the like.

[0054] The above application can be an embedded application installed in the device (i.e., the system application of the device), or a downloadable application. Among them, the embedded application is an application provided as part of the implementation of the device (such as a mobile phone). The downloadable application is an application that can provide its own Internet Protocol Multimedia Subsystem (IMS) connection. The downloadable application can be an application pre-installed in the device or a third-party application that can be downloaded and installed in the device by the user.

[0055] Reference Figure 1 , taking the above image enhancement device as a mobile phone as an example, Figure 1 shows a hardware structure of the mobile phone 10. As Figure 1 shown, the mobile phone 10 may include a processor 110, an external memory interface 120, an internal memory 121, a touch screen 130, an antenna 140, etc. Among them, the touch screen 130 includes a display screen 131 and a touchpad 132.

[0056] The processor 110 may include one or more processing units. For example: the processor 110 may include an application processor (AP), a modem processor, an image signal processor (ISP), a graphics processing unit (GPU), a controller, a memory, a video codec, a digital signal processor (DSP), a baseband processor, and / or a neural-network processing unit (NPU), etc. Among them, different processing units may be independent devices or integrated in one or more processors.

[0057] Among them, the controller may be the nerve center and command center of the mobile phone 10. The controller can generate operation control signals according to the instruction operation code and timing signal to complete the control of fetching instructions and executing instructions.

[0058] A memory can also be provided in the processor 110 for storing instructions and data. In some embodiments, the memory in the processor 110 is a cache memory. This memory can hold the instructions or data that the processor 110 has just used or recycled. If the processor 110 needs to use the instruction or data again, it can directly call it from the memory. This avoids repeated accesses, reduces the waiting time of the processor 110, and thus improves the efficiency of the system.

[0059] In some embodiments, the processor 110 may include one or more interfaces. The interfaces may include an inter-integrated circuit (I2C) interface, an inter-integrated circuit sound (I2S) interface, a pulse code modulation (PCM) interface, a universal asynchronous receiver / transmitter (UART) interface, a mobile industry processor interface (MIPI), a general-purpose input / output (GPIO) interface, a subscriber identity module (SIM) interface, and / or a universal serial bus (USB) interface, etc.

[0060] It can be understood that the mobile phone 10 can adopt the different interface connection methods described above, or a combination of multiple interface connection methods.

[0061] It should be understood that the mobile phone 10 can implement the shooting function through an ISP, a camera, a video codec, a GPU, a touch screen 130, and an AP, etc. The ISP is used to process the data fed back by the camera. The camera is used to obtain static images or videos. An object generates an optical image through a lens and projects it onto a photosensitive element. The DSP is used to process digital signals and can process other digital signals in addition to digital image signals. The video codec is used to compress or decompress digital videos. As an example, the ISP, the camera, the video codec, the GPU, the touch screen 130, and the AP, etc., can be used to shoot the to-be-processed image described above.

[0062] Among them, the display screen 131 of the touch screen 130 is used to display images, videos, etc. As an example, the display screen 131 can be used to display the to-be-processed image and the enhanced to-be-processed image described above. The touchpad 132 of the touch screen 130 can be used to input user instructions, etc.

[0063] It should be understood that the mobile phone 10 realizes the display function through the GPU, the display screen 131, the AP, etc. The GPU is a microprocessor for image processing, which is connected to the display screen 131 and the application processor AP. The GPU is used to perform mathematical and geometric calculations for graphics rendering. The processor 110 may include one or more GPUs, which execute program instructions to generate or change display information.

[0064] It should be understood that the NPU is a neural-network (NN) computing processor. By referring to the structure of the biological neural network, for example, referring to the transmission mode between human brain neurons, it can quickly process the input information and can also continuously self-learn. Through the NPU, applications such as intelligent processing of the mobile phone 10 can be realized, such as image enhancement, etc.

[0065] The external memory interface 120 can be used to connect an external memory card, such as a Micro SD card, to expand the storage capacity of the mobile phone 10. The external memory card communicates with the processor 110 through the external memory interface 120 to implement the data storage function. For example, files such as images are saved in the external memory card.

[0066] The internal memory 121 can be used to store computer-executable program codes, and the executable program codes include instructions. The processor 110 executes various functional applications and data processing of the mobile phone 10 by running the instructions stored in the internal memory 121.

[0067] The antenna 140 transmits and receives electromagnetic wave signals. Each antenna in the mobile phone 10 can be used to cover a single or multiple communication frequency bands. Different antennas can also be multiplexed to improve the utilization rate of the antennas. For example, the antenna 1 can be multiplexed as a diversity antenna for a wireless local area network. In some other embodiments, the antenna can be used in combination with a tuning switch.

[0068] It can be understood that the structure illustrated in the embodiments of the present application does not constitute a specific limitation on the mobile phone 10. In some other embodiments of the present application, the mobile phone 10 may include more or fewer components than those illustrated, or combine certain components, or split certain components, or have different component arrangements. The illustrated components can be implemented in hardware, software, or a combination of software and hardware.

[0069] The embodiments of the present application also provide another image enhancement device, which is used to train a neural network to obtain an image enhancement model that can implement the above image enhancement method. The image enhancement device can be a server or any other computing device with the computing power to train a neural network.

[0070] Reference Figure 2 , Figure 2The figure shows a schematic diagram of the hardware structure of a server provided by an embodiment of the present application. As Figure 2 shown, the server 20 may include a processor 21, a memory 22, a communication interface 23, and a bus 24. Among them, the processor 21, the memory 22, and the communication interface 23 may be connected through the bus 24.

[0071] The processor 21 is the control center of the server 20, and may be a general-purpose central processing unit (CPU), or other general-purpose processors, etc. Among them, the general-purpose processor may be a microprocessor or any conventional processor, etc.

[0072] As an example, the processor 21 may include one or more CPUs, such as Figure 2 the CPUs 0 and CPU1 shown in

[0073] The memory 22 may be a read-only memory (ROM) or other types of static storage devices that can store static information and instructions, a random access memory (RAM) or other types of dynamic storage devices that can store information and instructions, or an electrically erasable programmable read-only memory (EEPROM), a disk storage medium, or other magnetic storage devices, or any other medium that can be used to carry or store desired program code in the form of instructions or data structures and can be accessed by a computer, but is not limited thereto.

[0074] In a possible implementation, the memory 22 may exist independently of the processor 21. The memory 22 may be connected to the processor 21 through the bus 24 for storing data, instructions, or program code. When the processor 21 calls and executes the instructions or program code stored in the memory 22, it can train an image enhancement model for implementing the image enhancement method provided by the embodiment of the present application.

[0075] In another possible implementation, the memory 22 may also be integrated with the processor 21.

[0076] The communication interface 23 is used to connect the server 20 to other devices (such as terminals, etc.) through a communication network. The communication network may be an Ethernet, a radio access network (RAN), a wireless local area network (WLAN), etc. The communication interface 23 may include a receiving unit for receiving data and a sending unit for sending data.

[0077] The bus 24 can be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, an Extended Industry Standard Architecture (EISA) bus, or the like. This bus can be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Figure 2 only a thick line is used to represent it in the figure, but it does not mean that there is only one bus or one type of bus.

[0078] It should be noted that Figure 2 the structure shown in the figure does not constitute a limitation on the control device. Except Figure 2 for the components shown, the server 20 may include more or fewer components than shown in the figure, or combine some components, or have different component arrangements.

[0079] Next, with reference to the accompanying drawings, the image enhancement method provided by the embodiments of the present application will be described.

[0080] Refer to Figure 3 , Figure 3 which shows a schematic flowchart of the image enhancement method provided by the embodiments of the present application. This method is applied to an image enhancement device, which can be a terminal or Figure 2 the server shown in the figure, and this is not limited. This method may include the following steps:

[0081] S101. The image enhancement device acquires an image to be processed.

[0082] Among them, the image to be processed can be an image taken under low light conditions, or an image taken under other shooting conditions and requiring enhancement, and this is not limited in the embodiments of the present application.

[0083] Optionally, the image enhancement device can acquire the image to be processed from a local picture library. Among them, the pictures in the picture library include pictures pre-shot and saved, pictures downloaded from the network, pictures transmitted via Bluetooth, pictures sent by social software, and video screenshots in videos, etc., and this is not limited.

[0084] As an example, the user can select and load the image to be processed from the local picture library through a touch operation on the touch screen of the image enhancement device (such as Figure 1 the touch screen 130 shown in the figure), or through the voice interaction module of the image enhancement device. Correspondingly, the image enhancement device can acquire the image to be processed.

[0085] Optionally, the image enhancement device can also capture pictures in real time and use the pictures obtained by real-time capture as the image to be processed.

[0086] Of course, the image enhancement device can also obtain the image to be processed in any other way, and the embodiments of the present application do not make specific limitations on this.

[0087] S102. The image enhancement device obtains the low-frequency feature map of the above-mentioned image to be processed through the first neural network.

[0088] Among them, the first neural network can be an attention to context encoding (ACE) network.

[0089] Optionally, the image enhancement device can obtain the low-frequency feature map of the image to be processed through the ACE network.

[0090] Next, with reference to Figure 4 the network structure diagram of the ACE network 40 shown in

[0091] Step 1. The image enhancement device uses the ACE network 40 to obtain the first feature map 42 of the image to be processed 41.

[0092] Among them, the first feature map 42 is used to determine the first low-frequency feature map 43 of the image to be processed.

[0093] Specifically, the ACE network 40 can use two convolutional kernels to perform convolutional operations with the image to be processed 41 respectively to obtain features Figure 1 and features Figure 2 .

[0094] Among them, the above two convolutional kernels can be convolutional kernels with two different receptive fields (or called fields of view (FOV)), such as Figure 4 the convolutional kernel 1 and the convolutional kernel 2 shown in

[0095] Exemplarily, the above convolutional kernel 1 can be a convolutional kernel with a size of 3×3 and a dilation rate of 2, that is, the receptive field of the convolutional kernel 1 is 5. The above convolutional kernel 2 can be a convolutional kernel with a size of 1×1, that is, the receptive field of the convolutional kernel 2 is 1. In this way, the ACE network 40 can perform dilated convolutional operations with the convolutional kernel 1 and the image to be processed 41 to obtain features Figure 1 . The ACE network 40 can perform ordinary convolutional operations with the convolutional kernel 2 and the image to be processed 41 to obtain features Figure 2 .

[0096] Exemplarily, the above-mentioned convolutional kernel 1 can be a convolutional kernel with a size of 5×5, that is, the receptive field of convolutional kernel 1 is 5. The above-mentioned convolutional kernel 2 can be a convolutional kernel with a size of 1×1, that is, the receptive field of convolutional kernel 2 is 1. In this way, the ACE network 40 can perform ordinary convolution operations on the image to be processed 41 using convolutional kernel 1 to obtain features Figure 1 . The ACE network 40 can perform ordinary convolution operations on the image to be processed 41 using convolutional kernel 2 to obtain features Figure 2 .

[0097] It can be understood that the ACE network 40 can, according to the sizes of convolutional kernel 1 and convolutional kernel 2, perform a padding operation on the image to be processed 41 before performing convolution operations on the image to be processed using convolutional kernel 1 and convolutional kernel 2 respectively, so that the size of the feature map output after the convolution operation (such as the above-mentioned features Figure 1 and features Figure 2 ) is the same as the size of the image to be processed.

[0098] It should be understood that if the size of the convolutional kernel is 1×1, there is no need to pad the image to be processed.

[0099] Next, the ACE network 40 subtracts the pixel values of the corresponding pixels in the features Figure 1 and features Figure 2 to obtain the contrast bit perception feature map C for determining the high-frequency feature map of the image to be processed a .

[0100] Among them, the corresponding pixels in the features Figure 1 and features Figure 2 refer to the pixels located at the same position in the features Figure 1 and features Figure 2 . In this case, subtracting the pixel values of the corresponding pixels in the features Figure 1 and features Figure 2 means subtracting the pixel values of the pixels located at the same position in the features Figure 1 and features Figure 2 .

[0101] As an example, the ACE network 40 can subtract the pixel value of pixel 1 in the feature Figure 1 from the pixel value of pixel 2 corresponding to pixel 1 in the feature Figure 2 to obtain the contrast bit perception feature map C for determining the high-frequency feature map of the image to be processed a . Among them, pixel 1 and pixel 2 are the pixels located at the same position in the features Figure 1 and features Figure 2 .

[0102] Refer to Figure 5 , Figure 5Exemplarily, pixels at the same positions in feature map 51 and feature map 52 are shown. As Figure 5 shown, pixel Z511 in feature map 51 and Z521 in feature map 52 are pixels at the same position. Similarly, pixel Z515 in feature map 51 and Z525 in feature map 52 are pixels at the same position, pixel Z519 in feature map 51 and Z529 in feature map 52 are pixels at the same position, and so on.

[0103] Then, ACE network 40 negates the pixel value of each pixel in the contrast bit perception feature map C a to obtain which is the first feature map 42 for determining the first low-frequency feature map 43 of the image to be processed.

[0104] Specifically, negating the pixel value of each pixel in C a can be that ACE network 40 subtracts the value of 1 from each pixel value in C a to obtain That is,

[0105] Step 2: The image enhancement device uses the first neural network (i.e., ACE network 40) to multiply the pixel values of the corresponding pixels of the first feature map 42 and the image to be processed 41 to obtain the first low-frequency feature map 43 of the image to be processed.

[0106] Among them, multiplying the pixel values of the corresponding pixels of the first feature map 42 and the image to be processed 41 means multiplying the pixel values of the pixels at the same positions in the first feature map 42 and the image to be processed 41.

[0107] Among them, the description of the pixels at the same positions in the first feature map 42 and the image to be processed 41 can refer to the description of the pixels at the same positions in feature Figure 1 and feature Figure 2 in step 2 above, which will not be elaborated here.

[0108] As an example, the image enhancement device can use ACE network 40 to multiply the pixel value of the first pixel in the first feature map 42 and the pixel value of the second pixel corresponding to the first pixel in the image to be processed 41, so as to obtain the first low-frequency feature map 43. Among them, the first pixel and the second pixel are pixels at the same positions in the first feature map 42 and the image to be processed 41.

[0109] Step 3 (optional): Based on the first low-frequency feature map 43, the image enhancement device uses ACE network 40 to obtain the second low-frequency feature map 47 with context information.

[0110] Among them, the context information is the context information of the pixels in the second low-frequency feature map 47.

[0111] Specifically, the ACE network 40 can first perform pooling on the first low-frequency feature map 43 to obtain the pooled low-frequency feature map 44.

[0112] Among them, the ACE network 40 can perform pooling on the first low-frequency feature map 43 in the form of maxpooling. Alternatively, the ACE network 40 can perform pooling on the first low-frequency feature map 43 in the form of averagepooling. Here, the embodiments of the present application do not make specific limitations on the specific pooling method.

[0113] Next, based on the pooled low-frequency feature map 44, the ACE network 40 obtains the low-frequency feature map 46 with context information through the non-local neural sub-network 45.

[0114] Among them, for any pixel in the pooled low-frequency feature map 44, the context information refers to the relationship information between this pixel and all the pixels other than this pixel in the pooled low-frequency feature map 44.

[0115] Optionally, if the pooled low-frequency feature map 44 is regarded as a two-dimensional matrix block, in this way, the non-local neural sub-network 45 can transpose the low-frequency feature map 44 to obtain the transposed low-frequency feature map 440. Then, the non-local neural sub-network 45 performs a convolution operation on the low-frequency feature map 44 and the transposed low-frequency feature map 440 to obtain the low-frequency feature map M. Then, the non-local neural sub-network 45 performs a convolution operation on the low-frequency feature map M and the transposed low-frequency feature map 440 to obtain the low-frequency feature map 46 with context information.

[0116] It is easy to understand that the process of the ACE network 40 obtaining the low-frequency feature map 46 with context information through the non-local neural sub-network 45 based on the pooled low-frequency feature map 44 can be understood as a process of further learning the features in the pooled low-frequency feature map 44.

[0117] Then, the ACE network 40 performs unpooling on the low-frequency feature map 46 to obtain the second low-frequency feature map 47.

[0118] Among them, the method by which the ACE network 40 performs unpooling on the low-frequency feature map 46 is the reverse of the method by which the ACE network 40 performs pooling on the first low-frequency feature map 43.

[0119] Here, the ACE network 40 performs de-pooling on the low-frequency feature map 46. Specifically, according to the size of the first low-frequency feature map 43, the ACE network 40 predicts the adjacent pixels of each pixel in the low-frequency feature map 46, so as to obtain the second low-frequency feature map 47.

[0120] Step 4 (optional): The image enhancement device uses the ACE network 40 to fuse the first low-frequency feature map 43 and the second low-frequency feature map 47 to obtain the target low-frequency feature map 48.

[0121] Optionally, the ACE network 40 can add the pixel values of the corresponding pixels of the first low-frequency feature map 43 and the second low-frequency feature map 47, so as to obtain the target low-frequency feature map, which is the low-frequency feature map of the image to be processed described in the embodiments of the present application.

[0122] Among them, adding the pixel values of the corresponding pixels of the first low-frequency feature map 43 and the second low-frequency feature map 47 means adding the pixel values of the pixels located at the same position in the first low-frequency feature map 43 and the second low-frequency feature map 47.

[0123] Among them, the description of the pixels located at the same position in the first low-frequency feature map 43 and the second low-frequency feature map 47 can refer to the description of the pixels located at the same position in the features Figure 1 and features Figure 2 in step 2 above, which will not be elaborated here.

[0124] As an example, the ACE network 40 can add the pixel value of pixel 1 in the first low-frequency feature map 43 and the pixel value of pixel 2 corresponding to this pixel 1 in the second low-frequency feature map 47, so as to obtain the target low-frequency feature map. Among them, pixel 1 and pixel 2 are the pixels located at the same position in the first low-frequency feature map 43 and the second low-frequency feature map 47.

[0125] Optionally, in the embodiments of the present application, the first low-frequency feature map 43 can also be directly used as the low-frequency feature map of the image to be processed described in the embodiments of the present application, and no specific limitation is made in this regard. It should be understood that in this case, the image enhancement device does not need to perform the above steps 3 and 4.

[0126] It should be understood that since most of the noise in the image is high-frequency noise, when determining the low-frequency feature map of the image to be processed as described above, most of the high-frequency information of the image to be processed is filtered out, that is, the above method realizes noise filtering of the image to be processed.

[0127] S103: The image enhancement device determines the first image based on the above low-frequency feature map.

[0128] Among them, the first image includes the basic information of the to-be-processed image to be reconstructed, and the basic information includes the contour information of the to-be-processed image, etc.

[0129] Based on the above-mentioned low-frequency feature map, the image enhancement device can determine the first image through the following steps:

[0130] S1031. Based on the above-mentioned low-frequency feature map, the image enhancement device uses a first reconstruction network (corresponding to the third neural network in the embodiment of the present application) to reconstruct the basic information of the to-be-processed image, and obtains a basic image (corresponding to the third image in the embodiment of the present application).

[0131] Among them, the first reconstruction network can be a U-shaped network (Unet network) combined with a first cross-domain transformation (CDT) module.

[0132] Here, the Unet network includes at least one layer of downsampling convolutional network (i.e., a network with a stride of 2) and at least one layer of deconvolution network. Among them, the first CDT module is applied in combination with the at least one layer of deconvolution network, and each layer of deconvolution network is connected to a first CDT module.

[0133] As an example, as Figure 6 shown, Figure 6 exemplarily shows the structure of the first reconstruction network 60. As Figure 6 shown, the first reconstruction network 60 includes four layers of convolutional networks, namely convolutional network 601, convolutional network 602, convolutional network 603, and convolutional network 604. Among them, convolutional network 601 is the first layer of convolutional network, and its input is the low-frequency feature map obtained in the above step S102, for example Figure 4 shown low-frequency feature map 48. Convolutional network 604 is the fourth layer of convolutional network, and its output is feature map 605. It can be understood that during the convolution process of convolutional networks 601-604, every time a layer of downsampling convolutional network is passed through, the size of the feature map output by this layer of convolutional network can be 1 / 2 of the size of the feature map input to this layer of convolutional network.

[0134] As Figure 6As shown, the first reconstruction network 60 further includes four deconvolution networks, namely deconvolution network 6011, deconvolution network 6021, deconvolution network 6031, and deconvolution network 6041. Among them, the deconvolution networks and the convolution networks correspond one by one. For example, deconvolution network 6011 corresponds to convolution network 601, that is, the size of the feature map output by deconvolution network 6011 is the same as the size of the input feature map of convolution network 601. Among them, the input of deconvolution network 6041 is the feature map 605 output by convolution network 604. It can be understood that during the deconvolution process of deconvolution networks 6011 - 6014, after passing through each deconvolution network layer, the size of the feature map output by this layer of deconvolution network can be twice the size of the feature map input to this layer of deconvolution network.

[0135] As Figure 6 shown, the first reconstruction network 60 further includes four first CDT modules, namely CDT module 6042, CDT module 6032, CDT module 6022, and CDT module 6012. Among them, CDT module 6042 is connected to deconvolution network 6041, and the input feature map of CDT module 6042 is the output feature map of deconvolution network 6041. At this time, CDT module 6042 and deconvolution network 6041 are in the same network layer. CDT module 6042 is also connected to deconvolution network 6031, and the output feature map of CDT module 6042 is the input feature map of deconvolution network 6031. Among them, the connection situations of CDT module 6032, CDT module 6022, and CDT module 6012 with the deconvolution networks are similar to the connection situation of CDT module 6042 with the deconvolution network, and will not be elaborated here.

[0136] Among them, the feature map 62 output by CDT module 6012 is the basic image reconstructed by the first reconstruction network 60.

[0137] In this way, the image enhancement device can reconstruct the basic information of the image to be processed based on the low-frequency feature map of the image to be processed by using the first reconstruction network as Figure 6 shown, so as to obtain the basic image.

[0138] It should be understood that if the image to be processed is an image taken under low-light conditions, then the brightness and / or contrast of the basic image reconstructed by using the first reconstruction network is higher than that of the image to be processed, that is, the brightness and / or contrast of this basic image is the same (or similar) to the brightness and / or contrast of an image taken under natural daylight.

[0139] Next, the CDT module in the above first reconstruction network will be described in detail.

[0140] Refer to Figure 7 , Figure 7The structural schematic diagram of the first CDT module 70 is shown. The first CDT module 70 can be any one of the CDT modules in Figure 6 . As shown in Figure 7 , the input of the first CDT module 70 is the feature map 71. If the first CDT module 70 is the CDT module 6042 in Figure 6 , then the feature map 71 can be the feature map output by the deconvolution network 6041 in Figure 6 . Similarly, if the first CDT module 70 is the CDT module 6012 in Figure 6 , then the feature map 71 can be the feature map output by the deconvolution network 6011 in Figure 6 . No specific limitation is made thereto.

[0141] As shown in Figure 7 , the feature map 71 can be convolved with the convolution kernels 1 and 2 respectively, and after taking the difference of the operation results and then taking the inverse, the feature map 72 (referred to as the third feature map in the embodiments of the present application) can be obtained. The feature map 72 is used to represent the contrast perception feature of the feature map 71. Then, the feature map 72 is multiplied by the corresponding pixels of the feature map 71, and the feature map 73 is obtained. The feature map 73 is the low-frequency feature map of the feature map 71. Among them, for the process of obtaining the feature map 72 from the feature map 71 and further obtaining the feature map 73, reference can be made to the process of obtaining the first feature map from the image to be processed 41 and further obtaining the first low-frequency feature map 43 in Figure 4 , which will not be elaborated here.

[0142] Next, the first CDT module 70 takes the feature map 73 and the feature map 71 as the overall 2-channel feature map 74, and determines the global feature v of the 2-channel feature map 74. The global feature v can be the average value of all pixels in the 2-channel feature map 74. Then, the first CDT module 70 multiplies the global feature v and the 2-channel feature map 74 point by point to obtain the 2-channel feature map 75. The 2-channel feature map 75 is the feature map output by the CDT module 70. It can be understood that the global feature in the CDT module can globally adjust the 2-channel feature map 74, so that the output 2-channel feature map 75 is closer to the real one.

[0143] It should be understood that when the first CDT module 70 is the CDT module 6012 in Figure 6 , the output 2-channel feature map 75 thereof is the feature map 62 shown in Figure 6 .

[0144] From the above description, it can be seen that the CDT module includes a process of obtaining the low-frequency feature map of the input feature map. This process can filter out most of the high-frequency information in the feature map after deconvolution, so as to further filter out the noise of the image to be processed.

[0145] It should be understood that the Unet network itself also has the function of filtering noise. However, when the deconvolution network in the Unet network reconstructs an image, some high-frequency information will be introduced. Therefore, by combining the CDT module in the deconvolution network of the Unet network, the introduced high-frequency information can be filtered out.

[0146] S1032 (optional), the image enhancement device enhances the brightness and / or contrast of the above-mentioned basic image through a constant α to obtain a basic image with enhanced brightness and / or contrast (corresponding to the fourth image in the embodiments of the present application).

[0147] Among them, the constant α can be a constant preset by the image enhancement device or a constant output by the above-mentioned first reconstruction network when outputting the basic image. The embodiments of the present application do not make any limitations in this regard.

[0148] In a possible implementation manner, the image enhancement device can directly enhance the brightness and / or contrast of the above-mentioned basic image, that is, the image enhancement device can directly perform a dot product of the constant α and the above-mentioned basic image to obtain the fourth image.

[0149] In another possible implementation manner, the image enhancement device first uses a preset neural network to learn the above-mentioned basic image to obtain a fourth feature map. The fourth feature map represents the color and / or contrast information of the basic image. Among them, the embodiments of the present application do not specifically limit the structure and parameters of the preset neural network, and only require that the size of the fourth feature map output by the preset neural network is the same as the size of the image to be processed.

[0150] Next, the image enhancement device normalizes (sigmoid) the pixel values in the fourth feature map to obtain a fifth feature map.

[0151] Among them, the pixel value of each pixel in the fourth feature map is usually any integer between 1 and 255. When the image enhancement device normalizes the pixel values in the fourth feature map, it usually scales the pixel values of 0-255 to values between 0 and 1. That is to say, the pixel values in the fifth feature map are values between 0 and 1.

[0152] Exemplarily, if the value of pixel 1 in the fourth feature map is 100, after normalization, the value of this pixel is 0.392 (i.e., 100 / 255). Another exemplarily, if the value of pixel 2 in the fourth feature map is 200, after normalization, the value of this pixel is 0.784 (i.e., 200 / 255).

[0153] Then, the image enhancement device can enhance the brightness and / or contrast of the fifth feature map, that is, the image enhancement device can perform a dot product of a constant α and the fifth feature map to obtain a fourth image.

[0154] S1033 (optional), the image enhancement device multiplies the base image with enhanced brightness and / or contrast and the pixel value of the corresponding pixel of the base image to obtain a first image.

[0155] Among them, multiplying the base image with enhanced brightness and / or contrast (i.e., the fourth image) and the pixel value of the corresponding pixel of the base image (i.e., the third image) means multiplying the pixel values of the pixels at the same position in the fourth image and the third image.

[0156] Among them, for the description of the pixels at the same position in the fourth image and the third image, reference can be made to the features in S102 above Figure 1 and the features Figure 2 in the description of the pixels at the same position, which will not be elaborated here.

[0157] As an example, the image enhancement device can multiply the pixel value of the sixth pixel in the third image and the pixel value of the seventh pixel corresponding to the sixth pixel in the fourth image to obtain a first image. Among them, the sixth pixel and the seventh pixel are the pixels at the same position in the third image and the fourth image.

[0158] It should be understood that the image enhancement device can also directly use the base image (i.e., the third image) obtained in step S1031 as the first image. In this case, steps S1032 and S1033 of this application embodiment do not need to be executed.

[0159] S104, the image enhancement device obtains the high-frequency feature map of the above-mentioned image to be processed through a second neural network.

[0160] Among them, the second neural network can be an ACE network. The network structure of the second neural network is the same as that of the first neural network, and the second neural network can share the network parameters of the first neural network. In this way, the second neural network can invert the pixel value of each pixel in the first feature map obtained by the first neural network, that is, obtain a second feature map for determining the high-frequency feature map of the image to be processed. In this way, the second neural network can determine the second feature map without a large number of convolution operations, thus saving computing power.

[0161] Among them, for the description of the second neural network inverting the pixel value of each pixel in the first feature map, reference can be made to the description of the ACE network 40 inverting the bit-aware feature map C a above in S102, which will not be elaborated here.

[0162] Next, in a possible implementation, the second neural network may multiply the second feature map by the pixel values of the corresponding pixels of the first image above to obtain a first high-frequency feature map.

[0163] Among them, multiplying the second feature map by the pixel values of the corresponding pixels of the first image means multiplying the pixel values of the pixels at the same positions in the second feature map and the first image.

[0164] Among them, the description of the pixels at the same positions in the second feature map and the first image can refer to the features in S102 above Figure 1 and the features Figure 2 The description of the pixels at the same positions therein will not be elaborated here.

[0165] As an example, the image enhancement device may multiply the pixel value of the third pixel in the second feature map by the pixel value of the fourth pixel corresponding to the third pixel in the first image, so as to obtain a first high-frequency feature map. Among them, the third pixel and the fourth pixel are the pixels at the same positions in the second feature map and the first image.

[0166] In another possible implementation, the second neural network may multiply the second feature map by the pixel values of the corresponding pixels of the image to be processed to obtain a first high-frequency feature map.

[0167] Among them, multiplying the second feature map by the pixel values of the corresponding pixels of the image to be processed means multiplying the pixel values of the pixels at the same positions in the second feature map and the image to be processed.

[0168] Among them, the description of the pixels at the same positions in the second feature map and the image to be processed can refer to the features in S102 above Figure 1 and the features Figure 2 The description of the pixels at the same positions therein will not be elaborated here.

[0169] As an example, the image enhancement device may multiply the pixel value of the third pixel in the second feature map by the pixel value of the fifth pixel corresponding to the third pixel in the image to be processed, so as to obtain a first high-frequency feature map. Among them, the third pixel and the fifth pixel are the pixels at the same positions in the second feature map and the image to be processed.

[0170] Since the network structures of the second neural network and the first neural network are the same, therefore, the process of the second neural network determining the second high-frequency feature map based on the first high-frequency feature map, and determining the target high-frequency feature map based on the first high-frequency feature map and the second high-frequency feature map can refer to the description of obtaining the target low-frequency feature map 48 in steps 3 and 4 in S102 above, which will not be elaborated here.

[0171] S105. The image enhancement device determines a second image based on the above high-frequency feature map.

[0172] Among them, the second image includes the detailed information of the to-be-processed image to be reconstructed, and the detailed information includes at least one of the edges or textures of the to-be-processed image, and no specific limitation is made thereto.

[0173] Specifically, the image enhancement device may use a second reconstruction network (corresponding to the fourth neural network in the embodiments of the present application) to reconstruct the detailed information of the to-be-processed image, so as to obtain a detailed image of the to-be-processed image to be reconstructed, that is, the above-mentioned second image.

[0174] Among them, the second reconstruction network may be a Unet network combined with a first CDT module. The network structure of the second reconstruction network is the same as the network structure of the first reconstruction network, that is, the network structures of the fourth neural network and the third neural network are the same.

[0175] For the network structure of the second reconstruction network, reference may be made to the description of the first reconstruction network structure in S1031 above, and details are not described here.

[0176] Among them, the deconvolution network in the second reconstruction network is combined with the second CDT module, and the specific description may refer to the description in S1031 above, and details are not described here.

[0177] It should be noted that the second CDT module may share the network parameters of the first CDT module. In this way, the second CDT module may invert the third feature map obtained by the first CDT module, that is, obtain a sixth feature map for determining the high-frequency feature map of the input feature map of the first CDT module. In this way, the second CDT module can determine the sixth feature map without a large number of convolution operations, thereby saving the computational amount of the system.

[0178] Among them, for the description of inverting the third feature map, reference may be made to the description of ACE network 40 inverting the contrast bit perception feature map C a in S102 above, and details are not described here.

[0179] It should be understood that when the second CDT module shares the third feature map obtained by the first CDT module, the network layer where the second CDT module is located is the same as the network where the first CDT module is located.

[0180] As an example, in the first reconstruction network, the third feature map of the first CDT module in the first network layer may be shared by the second CDT module in the first network layer in the second reconstruction network.

[0181] In addition, the second CDT module determines the global feature v, and the description of making global adjustment to the deconvolution-reconstructed image using the global feature v can be referred to the above text and will not be elaborated here.

[0182] Optionally, the image enhancement device can also perform a further convolution operation on the detail image output by the second reconstruction network, so as to obtain a detail image after further convolution processing. In this case, the image enhancement device uses the detail image after the further convolution operation as the second image.

[0183] Among them, the above deep learning can be implemented through a neural network. In the embodiments of the present application, the network structure and network parameters of the neural network are not specifically limited, as long as the size of the feature map output by the neural network is the same as the size of the image to be processed.

[0184] S106. The image enhancement device fuses the first image and the second image to obtain an enhanced image of the image to be processed.

[0185] Optionally, the image enhancement device can add the pixel values of the corresponding pixels of the first image and the second image to obtain the enhanced image of the image to be processed.

[0186] Among them, adding the pixel values of the corresponding pixels of the first image and the second image means adding the pixel values of the pixels located at the same position in the first image and the second image.

[0187] Among them, the description of the pixels located at the same position in the first image and the second image can be referred to the features in S102 above Figure 1 and features Figure 2 in the description of the pixels located at the same position, which will not be elaborated here.

[0188] As an example, the image enhancement device can add the pixel value of the eighth pixel in the first image and the pixel value of the ninth pixel corresponding to the eighth pixel in the second image, so as to obtain the enhanced image of the image to be processed. Among them, the eighth pixel and the ninth pixel are pixels located at the same position in the first image and the second image.

[0189] So far, for the image enhancement method provided by the present application, the low-frequency feature map of the image to be processed is used to denoise and reconstruct the basic image, and the high-frequency feature map of the image to be processed is used to reconstruct the detail image, and then the basic image and the detail image are fused to obtain the enhanced image of the image to be processed. Through this method, while enhancing the brightness and / or contrast of the image, the noise in the image to be processed can be effectively filtered out.

[0190] In addition, in the image enhancement method provided in the embodiments of the present application, the second neural network can share the first feature map of the first neural network, and the second CDT module can share the third feature map of the first CDT module. In this way, during the process of enhancing the image to be processed by the above method, the image enhancement device can reduce a large number of convolution operations, thereby saving computing power.

[0191] It should be understood that in practical applications, the image enhancement method provided in the embodiments of the present application can be implemented by directly executing the above steps S101-S106 through an image enhancement device, or can be implemented by pre-setting an image enhancement model capable of implementing the above method in the image enhancement device.

[0192] Next, the structure of the image enhancement model will be briefly described:

[0193] Reference Figure 8 , Figure 8 shows a schematic structural diagram of an image enhancement model. As Figure 8 shown, the image enhancement model 80 includes a first neural network module 81, a first reconstruction module 82, a second neural network module 83, a second reconstruction module 84, and a fusion module 85. Among them, the first neural network module 81 and the first reconstruction module 82 can be used as the neural network modules in the first stage of the image enhancement model 80, and the second neural network module 83 and the second reconstruction module 84 can be used as the neural network modules in the second stage of the image enhancement model 80.

[0194] Among them, the first neural network module 81 can include the first neural network described above and is used to implement the function of obtaining the low-frequency feature map of the image to be processed in step S102 above.

[0195] The first reconstruction module 82 can include a first reconstruction network sub-module 821. The first reconstruction network sub-module 821 can include the first reconstruction network described above and is used to implement the function of reconstructing the basic image (i.e., the third image) of the image to be processed in step S1031 above.

[0196] Optionally, the first reconstruction module 82 can further include an enhancement sub-module 822. The enhancement sub-module 822 can be used to implement the function of obtaining the fourth image after enhancing the brightness and / or contrast of the basic image in step S1032 above. It should be understood that the enhancement sub-module 822 can include the preset neural network described in step S1032 above, and the preset neural network is used to obtain the fourth feature map described above, and the fourth feature map can be used to obtain the fourth image. The first reconstruction module 82 is further used to implement the function of multiplying the basic image after enhancing the brightness and / or contrast and the pixel values of the corresponding pixels of the basic image to obtain the first image in step S1033 above.

[0197] The second neural network module 83 may include the second neural network described above and is used to implement the function of obtaining the high-frequency feature map of the image to be processed in step S104 above.

[0198] The second reconstruction module 84 may include the second reconstruction network described above and is used to implement the function of determining the second image based on the high-frequency feature map in step S105 above.

[0199] The fusion module 85 may be used to implement the function of fusing the first image output by the first reconstruction module 82 and the second image output by the second reconstruction network module 84 to obtain the enhanced image of the image to be processed in step S106 above.

[0200] Among them, the functions and beneficial effects realized by each module in the first neural network, the second neural network, the first reconstruction network, the second reconstruction network, and the image enhancement model 80 may refer to the descriptions in S101 - S106 above and will not be elaborated here.

[0201] It should be understood that the above image enhancement model can be pre-trained by an image enhancement device (such as Figure 2 the server shown), or can be pre-trained by any device with the ability to train a neural network model.

[0202] Next, taking the image enhancement device pre-training the above image enhancement model as an example, the method for the image enhancement device to train the image enhancement model will be described.

[0203] Refer to Figure 9 , Figure 9 which shows a schematic flowchart of the method for the image enhancement device to train the image enhancement model. The method may include the following steps:

[0204] S201. The image enhancement device obtains at least one training sample.

[0205] The description of the image enhancement device obtaining the training sample may refer to the description of obtaining the image to be processed in S101 above and will not be elaborated here.

[0206] Among them, any one of the at least one training sample includes a training image pair, and the training image pair includes a training image and a training target image. Among them, the training image can be used as the image to be enhanced, and the training target image can be used as the target image after the training image is enhanced.

[0207] Among them, the training image and the training target image can be standard RGB format images.

[0208] It should be understood that the sizes of the training image and the training target image in a pair of training images are usually the same. Herein, the embodiments of the present application do not specifically limit the sizes of the training image and the training target image in a pair of training images. Of course, if the sizes of the training image and the training target image are small (such as 512×384), the computing power required for the image enhancement device to train the image enhancement model can be saved.

[0209] It should be understood that in a pair of training images, the image content in the training image is the same as or similar to the image content in the training target image.

[0210] Exemplarily, if the pair of training images A includes the training image A and the training target image A, the training image A may be an image of scene A taken under low light conditions, and the training target image A may be an image of scene A taken under natural daylight conditions. Among them, the training image A and the training target image A are images of scene A taken at the same or similar shooting angles.

[0211] Optionally, to increase the number of training samples, the images (including the training image and the training target image) in the existing training samples can be randomly flipped horizontally or vertically. Exemplarily, after the images in the training sample A are flipped horizontally, a training sample B can be added.

[0212] S202. The image enhancement device trains an image enhancement model according to at least one of the above training samples.

[0213] Specifically, the image enhancement device can iteratively train a neural network according to at least one of the above training samples, so as to obtain the above image enhancement model.

[0214] Specifically, if at least one of the above training samples includes m (m is an integer greater than or equal to 1) training samples, the process of the image enhancement device using the m training samples to iteratively train the neural network to obtain the image enhancement model may include the following steps:

[0215] Step 1: The image enhancement device inputs the training image 1 in the training sample 1 among the m training samples into the initial image enhancement model.

[0216] Among them, the structure of the initial image enhancement model is as Figure 8 shown, including multiple neural networks, which will not be elaborated here.

[0217] After the image enhancement device inputs the training image 1 into the initial image enhancement model, the initial image enhancement model can output the target image 1 through the method described in S101-S106 above.

[0218] Then, the image enhancement device can calculate the loss function 1 based on the target image 1 and the training target image 1 in the training sample 1. It can be seen that the training image 1 and the training target image 1 belong to the same training image pair.

[0219] Optionally, the image enhancement device can also calculate the loss function 2 based on the target image 1 and the third image obtained during the process of enhancing the training image 1 by the initial image enhancement model. For the description of the third image obtained during the process of enhancing the training image 1 by the initial image enhancement model, reference can be made to the description of obtaining the third image in the above text, which will not be elaborated here.

[0220] Then, the image enhancement device feeds the loss function 1 and the loss function 2 back to the initial image enhancement model respectively to adjust the parameters of the neural network in the initial image enhancement model, so as to obtain the image enhancement model 2 with adjusted neural network parameters. Or, the image enhancement device fuses the loss function 1 and the loss function 2 into the loss function 0, and then feeds the loss function 0 back to the initial image enhancement model to adjust the parameters of the neural network in the initial image enhancement model, so as to obtain the image enhancement model 2 with adjusted neural network parameters.

[0221] Step 2: The image enhancement device inputs the training image 2 in the training sample 2 into the image enhancement model 2, and refers to the above step 1 to obtain the image enhancement model 3.

[0222] In this way, through multiple rounds of iterative training, when the number of iterative rounds reaches the preset threshold, or the loss function calculated by the image enhancement device is less than or equal to the preset threshold, the current image enhancement model is output as the target image enhancement model, that is, the image enhancement device has trained Figure 8 the described image enhancement model.

[0223] In this case, when the image enhancement device trains the target image enhancement model, it can release it as a dedicated image enhancement App. In this way, after the user installs / updates the image enhancement App, the user can enhance images through this App. Or, after the image enhancement device trains the target image enhancement model, it can be applied to the image processing App as an image enhancement function module of the App. In this way, when the user installs / updates the image processing App including the image enhancement model, the user can use the image enhancement model in the App to enhance images.

[0224] Exemplarily, take Figure 1In the mobile phone 10 shown, an image processing App including an image enhancement model is installed. Then, the user can click on the "Image Processing" App icon on the touch screen of the mobile phone 10 to open the image processing App. Next, the user can load the image to be processed on the display interface of the image processing App, and in the picture editing state, click the "Image Enhancement" button in the toolbar to enhance the image to be processed. At this time, the enhanced image to be processed can be displayed on the display interface of the image processing App.

[0225] In summary, the image enhancement method provided in this application can pre-train an image enhancement model, use the low-frequency feature map of the image to be processed obtained to denoise and reconstruct the basic image, and use the high-frequency feature map of the image to be processed obtained to reconstruct the detail image. Then, by fusing the basic image and the detail image, the enhanced image of the image to be processed is obtained. Through this method, while enhancing the brightness and / or contrast of the image, the noise in the image to be processed can be effectively filtered out.

[0226] In addition, in the above image enhancement model, the second neural network can share the first feature map of the first neural network, and the second CDT module can share the third feature map of the first CDT module. In this way, when enhancing the image through this image enhancement model, a large number of convolution operations are reduced, thus saving computing power.

[0227] The above mainly introduces the solution provided in the embodiments of this application from the perspective of the method. To implement the above functions, it includes the corresponding hardware structure and / or software module for executing each function. Those skilled in the art should easily realize that, combining the units and algorithm steps of each example described in the embodiments disclosed in this article, this application can be implemented in the form of hardware or a combination of hardware and computer software. Whether a certain function is executed in the way of hardware or computer software driving hardware depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of this application.

[0228] The embodiments of this application can divide the functional modules of the image enhancement device according to the above method examples. For example, each functional module can be divided corresponding to each function, or two or more functions can be integrated into one processing module. The above integrated module can be implemented in the form of hardware or in the form of a software functional module. It should be noted that the division of modules in the embodiments of this application is illustrative, only a logical function division, and there may be other division methods in actual implementation.

[0229] As Figure 10 shown, Figure 10The figure shows a schematic structural diagram of an image enhancement device 100 provided by an embodiment of the present application. The image enhancement device 100 can be used to execute the above-mentioned image enhancement method, for example, to execute Figure 3 the method shown. Among them, the image enhancement device 100 may include an acquisition unit 101, a determination unit 102, and a fusion unit 103.

[0230] The acquisition unit 101 is configured to obtain a low-frequency feature map of the image to be processed through a first neural network. The determination unit 102 is configured to determine a first image according to the low-frequency feature map obtained by the acquisition unit 101. Wherein, the first image includes the basic information of the reconstructed image to be processed, and the basic information includes the contour information of the image to be processed. Then, the acquisition unit 101 is further configured to obtain a high-frequency feature map of the image to be processed through a second neural network. The determination unit 102 is further configured to determine a second image according to the high-frequency feature map obtained by the acquisition unit 101. Wherein, the second image includes the detailed information of the reconstructed image to be processed, and the detailed information includes at least one of the edges or textures of the image to be processed. Then, the fusion unit 103 is configured to fuse the first image and the second image to obtain the enhanced image of the image to be processed.

[0231] As an example, in combination with Figure 3 , the acquisition unit 101 can be used to execute S102 and S104, the determination unit 102 can be used to execute S103 and S105, and the fusion unit 103 can be used to execute S106.

[0232] Optionally, the acquisition unit 101 is specifically configured to: use a first neural network to obtain a first feature map of the image to be processed; and, use the first neural network to multiply the pixel value of a first pixel in the first feature map by the pixel value of a second pixel corresponding to the first pixel in the image to be processed to obtain a low-frequency feature map of the image to be processed. The acquisition unit 101 is further specifically configured to: use a second neural network to invert the pixel value of each pixel in the first feature map to obtain a second feature map; and, use the second neural network to multiply the pixel value of a third pixel in the second feature map by the pixel value of a fourth pixel corresponding to the third pixel in the first image, or, use the second neural network to multiply the pixel value of the third pixel by the pixel value of a fifth pixel corresponding to the third pixel in the image to be processed to obtain a high-frequency feature map of the image to be processed. Wherein, the network structures of the second neural network and the first neural network are the same.

[0233] As an example, in combination with Figure 3 , the acquisition unit 101 can be used to execute S102 and S104.

[0234] Optionally, the above-mentioned "inversion" means subtracting 1 from the pixel value of each pixel in the first feature map.

[0235] Optionally, the determination unit 102 is specifically configured to reconstruct the basic information of the image to be processed based on the low-frequency feature map obtained by the acquisition unit 101 using a third neural network to obtain a third image; and enhance the color and / or contrast of the third image by a constant α to obtain a fourth image; and multiply the pixel value of the sixth pixel in the third image by the pixel value of the seventh pixel corresponding to the sixth pixel in the fourth image to obtain a first image.

[0236] As an example, in combination with Figure 3 , the determination unit 102 can be used to execute S1031 to S1033.

[0237] Optionally, the above constant α is a preset constant, or the above constant α is obtained through a third neural network.

[0238] Optionally, the determination unit 102 is further specifically configured to reconstruct the detailed information of the image to be processed based on the high-frequency feature map obtained by the acquisition unit 101 using a fourth neural network to obtain a second image; wherein, the network structure of the fourth neural network is the same as that of the above-mentioned third neural network. The feature map used to obtain the high-frequency feature map in the fourth neural network is obtained by inverting the pixel value of each pixel in the feature map used to obtain the low-frequency feature map in the above-mentioned third neural network.

[0239] As an example, in combination with Figure 3 , the determination unit 102 can be used to execute S105.

[0240] Optionally, the fusion unit 103 is specifically configured to add the pixel value of the eighth pixel in the first image to the pixel value of the ninth pixel corresponding to the eighth pixel in the second image to obtain the enhanced image of the image to be processed.

[0241] As an example, in combination with Figure 3 , the fusion unit 103 can be used to execute S106.

[0242] For the specific description of the above optional manner, reference may be made to the foregoing method embodiments, which will not be elaborated herein. In addition, the explanations and beneficial effects descriptions of any of the above-provided image enhancement devices 100 can refer to the corresponding method embodiments above, and will not be elaborated.

[0243] As an example, in combination with Figure 1 , the acquisition unit 101, the determination unit 122, and the fusion unit 103 in the image enhancement device 100 can be implemented by the processor 101 in Figure 1 executing the program code in the internal memory 121 in Figure 1 .

[0244] The embodiment of the present application further provides a chip system 110, such as Figure 11As shown, the chip system 110 includes at least one processor and at least one interface circuit. As an example, when the chip system 110 includes one processor and one interface circuit, the one processor may be Figure 11 the processor 111 shown in the solid line box in Figure 11 (or the processor 111 shown in the dashed line box), and the one interface circuit may be Figure 11 the interface circuit 112 shown in the solid line box in Figure 11 (or the interface circuit 112 shown in the dashed line box). When the chip system 110 includes two processors and two interface circuits, the two processors include

[0245] the processor 111 shown in the solid line box and the processor 111 shown in the dashed line box in

[0246] and the two interface circuits include

[0247] the interface circuit 112 shown in the solid line box and the interface circuit 112 shown in the dashed line box in

[0248] Figure 12 This is not limited.

[0249] The processor 111 and the interface circuit 112 can be interconnected by lines. For example, the interface circuit 112 can be used to receive signals (such as obtaining an image to be processed, etc.). Also for example, the interface circuit 112 can be used to send signals to other devices (such as the processor 111). Exemplarily, the interface circuit 112 can read the instructions stored in the memory and send the instructions to the processor 111. When the instructions are executed by the processor 111, the image enhancement device can execute each step in the above-mentioned embodiments. Of course, the chip system 110 may also include other discrete devices, and the embodiments of the present application do not make specific limitations on this. Figure 3 or Figure 9 described functions or partial functions. Therefore, for example, referring toFigure 3 One or more of the features in S101 - S106 may be carried by one or more instructions associated with the signal - bearing medium 120. Additionally, Figure 12 The program instructions in also describe example instructions.

[0250] In some examples, the signal - bearing medium 120 may include a computer - readable medium 121, such as but not limited to, a hard - disk drive, a compact disc (CD), a digital video disc (DVD), a digital tape, a memory, a read - only memory (ROM), or a random access memory (RAM), etc.

[0251] In some embodiments, the signal - bearing medium 120 may include a computer - recordable medium 122, such as but not limited to, a memory, a read / write (R / W) CD, an R / W DVD, etc.

[0252] In some embodiments, the signal - bearing medium 120 may include a communication medium 123, such as but not limited to, a digital and / or analog communication medium (e.g., an optical fiber cable, a waveguide, a wired communication link, a wireless communication link, etc.).

[0253] The signal - bearing medium 120 may be conveyed by a wireless form of the communication medium 123 (e.g., a wireless communication medium compliant with the IEEE 1202.11 standard or other transmission protocols). One or more program instructions may be, for example, computer - executable instructions or logic implementation instructions.

[0254] In some examples, such as for Figure 3 or Figure 9 The image - enhancement device described may be configured to provide various operations, functions, or actions in response to one or more program instructions via the computer - readable medium 121, the computer - recordable medium 122, and / or the communication medium 123.

[0255] It should be understood that the arrangements described herein are for illustrative purposes only. Thus, those skilled in the art will understand that other arrangements and other elements (e.g., machines, interfaces, functions, sequences, and function groups, etc.) can be used instead, and some elements may be omitted altogether depending on the desired results. Additionally, many of the elements described can be implemented as discrete or distributed components, or as functional entities that combine with other components in any suitable combination and location.

[0256] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using a software program, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer instructions are executed on a computer, the processes or functions according to the embodiments of the present application are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center in a wired manner (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or a wireless manner (such as infrared, wireless, microwave, etc.). The computer-readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server or a data center that includes one or more integrated media. The available medium can be a magnetic medium (such as a floppy disk, hard disk, magnetic tape), an optical medium (such as a DVD), or a semiconductor medium (such as a solid state disk (SSD)), etc.

[0257] As described above, the above are only specific embodiments of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention can easily think of changes or substitutions, which should all be covered by the protection scope of the present invention. Therefore, the protection scope of the present invention should be subject to the protection scope of the claims.

Claims

1. An image enhancement method, characterized in that, the method includes: obtaining a low-frequency feature map of the image to be processed through a first neural network; determining a first image according to the low-frequency feature map, the first image including the basic information of the reconstructed image to be processed, and the basic information including the contour information of the image to be processed; obtaining a high-frequency feature map of the image to be processed through a second neural network; determining a second image according to the high-frequency feature map, the second image including the detailed information of the reconstructed image to be processed, and the detailed information including at least one of the edges or textures of the image to be processed; fusing the first image and the second image to obtain an enhanced image.

2. The method according to claim 1, characterized in that, the step of obtaining a low-frequency feature map of the image to be processed through a first neural network includes: using the first neural network to obtain a first feature map of the image to be processed; and multiplying the pixel value of a first pixel in the first feature map by the pixel value of a second pixel corresponding to the first pixel in the image to be processed using the first neural network to obtain the low-frequency feature map of the image to be processed; the step of obtaining a high-frequency feature map of the image to be processed through a second neural network includes: using the second neural network to invert the pixel value of each pixel in the first feature map to obtain a second feature map; and multiplying the pixel value of a third pixel in the second feature map by the pixel value of a fourth pixel corresponding to the third pixel in the first image, or multiplying the pixel value of the third pixel by the pixel value of a fifth pixel corresponding to the third pixel in the image to be processed using the second neural network to obtain the high-frequency feature map of the image to be processed; wherein, the network structures of the second neural network and the first neural network are the same.

3. The method according to claim 2, characterized in that, the inversion means subtracting 1 from the pixel value of each pixel in the first feature map.

4. The method according to any one of claims 1-3, characterized in that, the step of determining a first image according to the low-frequency feature map includes: reconstructing the basic information of the image to be processed according to the low-frequency feature map using a third neural network to obtain a third image; enhancing at least one of the color or contrast of the third image by a constant α to obtain a fourth image; multiplying the pixel value of a sixth pixel in the third image by the pixel value of a seventh pixel corresponding to the sixth pixel in the fourth image to obtain the first image.

5. The method according to claim 4, characterized in that, the constant α is a preset constant, or the constant α is obtained through the third neural network.

6. The method according to claim 4, characterized in that, the step of determining a second image according to the high-frequency feature map includes: reconstructing the detailed information of the image to be processed according to the high-frequency feature map using a fourth neural network to obtain a second image; Among them, the network structures of the fourth neural network and the third neural network are the same; the feature map used to obtain the high-frequency feature map in the fourth neural network is obtained by taking the inverse of the pixel value of each pixel in the feature map used to obtain the low-frequency feature map in the third neural network.

7. The method according to any one of claims 1-3, wherein, the step of fusing the first image and the second image to obtain an enhanced image includes: adding the pixel value of the eighth pixel in the first image and the pixel value of the ninth pixel corresponding to the eighth pixel in the second image to obtain the enhanced image.

8. An image enhancement device, wherein, the device includes: an acquisition unit, configured to obtain a low-frequency feature map of an image to be processed through a first neural network; a determination unit, configured to determine a first image according to the low-frequency feature map, where the first image includes the basic information of the reconstructed image to be processed, and the basic information includes the contour information of the image to be processed; the acquisition unit is further configured to obtain a high-frequency feature map of the image to be processed through a second neural network; the determination unit is further configured to determine a second image according to the high-frequency feature map, where the second image includes the detailed information of the reconstructed image to be processed, and the detailed information includes at least one of the edges or textures of the image to be processed; a fusion unit, configured to fuse the first image and the second image to obtain an enhanced image.

9. The image enhancement device according to claim 8, wherein, the acquisition unit is specifically configured to: use the first neural network to obtain a first feature map of the image to be processed; and use the first neural network to multiply the pixel value of the first pixel in the first feature map by the pixel value of the second pixel corresponding to the first pixel in the image to be processed to obtain the low-frequency feature map of the image to be processed; and use the second neural network to invert the pixel value of each pixel in the first feature map to obtain a second feature map; and use the second neural network to multiply the pixel value of the third pixel in the second feature map by the pixel value of the fourth pixel corresponding to the third pixel in the first image, or use the second neural network to multiply the pixel value of the third pixel by the pixel value of the fifth pixel corresponding to the third pixel in the image to be processed to obtain the high-frequency feature map of the image to be processed; wherein, the network structures of the second neural network and the first neural network are the same.

10. The image enhancement device according to claim 9, wherein, the inversion means subtracting 1 from the pixel value of each pixel in the first feature map.

11. The image enhancement device according to any one of claims 8-10, wherein, The determining unit is specifically configured to reconstruct the basic information of the image to be processed based on the low-frequency feature map by using a third neural network to obtain a third image; and enhance the color and / or contrast of the third image by a constant α to obtain a fourth image; and multiply the pixel value of the sixth pixel in the third image by the pixel value of the seventh pixel corresponding to the sixth pixel in the fourth image to obtain the first image.

12. The image enhancement device according to claim 11, wherein, the constant α is a preset constant, or the constant α is obtained by the third neural network.

13. The image enhancement device according to claim 11, wherein, the determining unit is further specifically configured to reconstruct the detailed information of the image to be processed based on the high-frequency feature map by using a fourth neural network to obtain a second image; wherein, the network structures of the fourth neural network and the third neural network are the same; the feature map for obtaining the high-frequency feature map in the fourth neural network is obtained by inverting the pixel value of each pixel in the feature map for obtaining the low-frequency feature map in the third neural network.

14. The image enhancement device according to any one of claims 8-10, wherein, the fusion unit is specifically configured to add the pixel value of the eighth pixel in the first image to the pixel value of the ninth pixel corresponding to the eighth pixel in the second image to obtain the enhanced image.

15. An image enhancement device, wherein, the device includes: a memory and one or more processors, the memory is used to store computer instructions, and the processor is used to call the computer instructions to execute the image enhancement method according to any one of claims 1-7.

16. A computer-readable storage medium, wherein, a computer program is stored on the computer-readable storage medium, and when the computer program runs on a computer, the computer is caused to execute the image enhancement method according to any one of claims 1-7.

Citation Information

Patent Citations

  • Image enhancement processing method and device

    CN108305236A

  • Image processing method and device and terminal equipment

    CN111429371A