Image restoration method, device and computer storage medium

By combining the spatial and photosensitive characteristics of RAW images for feature extraction, an image restoration network is guided to restore RAW images, solving the problems of detail loss and artifacts in existing technologies and achieving high-quality image restoration.

CN116228554BActive Publication Date: 2026-01-02HUAWEI TECH CO LTD +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211632316.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-12-19
Publication Date
2026-01-02
Estimated Expiration
2042-12-19

AI Technical Summary

Technical Problem

Existing technologies are prone to detail loss and severe artifacts during image restoration, failing to effectively improve the quality of RAW images.

Method used

By combining the spatial and photosensitive characteristics of RAW images, feature extraction is performed using photosensitive and spatial guide maps to guide the image restoration network in image restoration, thereby enhancing details and textures and reducing artifacts.

Benefits of technology

It effectively enhances the details and textures of restored RAW images, reduces artifacts, and ensures image quality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116228554B_ABST
    Figure CN116228554B_ABST
Patent Text Reader

Abstract

The application provides an image restoration method, device and computer storage medium. In an embodiment, the method comprises: obtaining a RAW image obtained by an image sensor when shooting a target scene; wherein the RAW image comprises N pixel channels; determining a light sensing guide map corresponding to the RAW image, wherein the light sensing guide map indicates brightness information of the RAW image; determining a spatial guide map corresponding to the RAW image; wherein the spatial guide map indicates a positional relationship of the N pixel channels; performing feature extraction to determine a guide feature of an image restoration network based on the light sensing guide map and the spatial guide map; and performing image restoration on the RAW image based on the guide feature and the image restoration network to obtain a restored image corresponding to the RAW image. The spatial characteristics and the light sensing characteristics of the RAW image are combined to guide the image restoration network to restore the RAW image, details and textures of the restored RAW image are improved, artifacts are reduced, and the image quality is ensured.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of image processing, and particularly relates to an image recovery method and device and computer storage medium. BACKGROUND

[0002] In the field of computer vision, photographing has become one of the most commonly used functions of various mobile terminal devices such as mobile phones, tablet computers, smart glasses and wearable devices. At present, the image obtained through an image sensor (such as a camera) is an unprocessed RAW image. However, due to the performance and power consumption of the image sensor, the quality recovery of the RAW image is a very urgent requirement, so as to improve the visual effect and help improve the accuracy of subsequent visual tasks such as target detection and scene segmentation.

[0003] At present, the channel with high sensitivity and high signal-to-noise ratio is first recovered to obtain a partial channel image after recovery. Then, the RAW image is further fused and processed according to the channel image after recovery, and finally a recovered image is obtained.

[0004] However, the above method is prone to cause loss of details and serious artifacts, and therefore, there is an urgent need for an image recovery method capable of preserving details.

[0005] The information disclosed in this Background section is only for the purpose of increasing the understanding of the general background of the application, and should not be taken as admission that the information forms the prior art base for the present application. SUMMARY

[0006] The embodiments of the present application provide an image recovery method, device and computer storage medium, which recover the RAW image by jointly guiding the image recovery network through the spatial characteristics and photosensitive characteristics of the RAW image, improve the details and texture of the RAW image after recovery, reduce artifacts, and ensure the image quality.

[0007] In a first aspect, an embodiment of the present application provides an image restoration method, comprising: obtaining a RAW image of a target scene captured by an image sensor; wherein the image sensor collects the RAW image through a front color filter matrix, the color filter matrix is arranged in target pixel units, each target pixel unit is arranged in N target pixels, the target pixels at the same positions of each target pixel unit in the color filter matrix form a pixel channel, and a total of N pixel channels are formed; determining a light sensing guide map corresponding to the RAW image, wherein the light sensing guide map indicates brightness information of the RAW image; determining a spatial guide map corresponding to the RAW image; wherein the spatial guide map indicates a positional relationship of the N pixel channels; based on the light sensing guide map and the spatial guide map, performing feature extraction to determine a guide feature of an image restoration network; and based on the guide feature, performing image restoration on the RAW image according to the image restoration network to obtain a restored image corresponding to the RAW image.

[0008] In the present solution, the guide feature considering the spatial characteristics and light sensing characteristics of the RAW image is used to assist the image restoration network to restore the RAW image, so that the image restoration network extracts information meeting the spatial characteristics and light sensing characteristics, improves the details and textures of the restored RAW image, reduces artifacts, and ensures the image quality.

[0009] In a possible implementation, the guide feature indicates the extracted light sensing information, and the feature obtained by the information of the positional relationship of the N pixel channels in the RAW image can reflect the light sensing difference of the N pixel channels and the difference in pixel arrangement of the target pixels of the N pixel channels. The pixel arrangement indicates the distribution of the target pixels belonging to the same pixel channel around the target pixel.

[0010] In a possible implementation, the image restoration network considers the light sensing difference of the channels of the N target pixels of the RAW image, and uses different processing to realize image restoration on the target pixels of the N pixel channels with different pixel arrangements. The pixel arrangement indicates the distribution of the target pixels belonging to the same pixel channel around the target pixel.

[0011] In a possible implementation, based on the guide feature, the image restoration on the RAW image according to the image restoration network comprises: based on the guide feature, processing at least part of the features extracted by the image restoration network from the RAW image.

[0012] Here, based on the guide feature, at least part of the features extracted by the image restoration network can be corrected, so that the corrected features have spatial characteristics and light sensing characteristics, thereby better restoring the details and textures of the RAW image and reducing artifacts.

[0013] In the scheme, at least part of the features extracted in the image restoration network are corrected by considering the guiding features of the spatial characteristics and the photosensitive characteristics of the RAW image, so that the image restoration network extracts information meeting the photosensitive characteristics and the spatial characteristics, improves the details and textures after the RAW image is restored, reduces the artifacts, and ensures the image quality.

[0014] In a possible implementation, the number of target layers in the image restoration network is M, the guiding features include guiding feature maps corresponding to m target layers in the M target layers, M is a positive integer greater than or equal to 2, and m is a positive integer less than or equal to M; the target layer is a convolution layer or a pooling layer; based on the guiding features, the RAW image is subjected to image restoration according to the image restoration network, including: in the process of image restoration of the RAW image according to the image restoration network, when a first target layer in the m target layers is passed, window data of a target window of the first target layer on an input graph is processed based on a guiding feature map corresponding to the first target layer; wherein the input graph is a graph input into the first target layer, and the first target layer is any target layer in the m target layers.

[0015] In the scheme, by considering the guiding features of the spatial characteristics and the photosensitive characteristics of the RAW image, on the one hand, different window data in the input graph sliding window process of the target window of the target layer is subjected to different guiding processing, and on the other hand, the features at any position of the image restoration network can be processed, so that the target layer extracts information meeting the photosensitive characteristics and the spatial characteristics, improves the details and textures after the RAW image is restored, reduces the artifacts, and ensures the image quality.

[0016] In an example, based on the guiding feature map corresponding to the first target layer, the window data in the input graph sliding window process of the target window of the first target layer is processed, including: determining window data of the target window of the first target layer; wherein the window data is data at a position of the target window on the input graph; determining first guiding data corresponding to the window data; wherein the first guiding data is data in the guiding feature map corresponding to the first target layer; and processing the window data based on the first guiding data.

[0017] In the scheme, by considering the guiding features of the spatial characteristics and the photosensitive characteristics of the RAW image, the window data is processed, so that information meeting the photosensitive characteristics and the spatial characteristics in the window data is extracted, the details and textures after the RAW image is restored are improved, the artifacts are reduced, and the image quality is ensured.

[0018] In a possible case, the first guiding data corresponding to the window data of the target window at different positions on the input graph is different.

[0019] In a possible case, the input image and the guide feature map corresponding to the first target layer are of the same plane size, and the first guide data is data in the guide feature map corresponding to the first target layer that is adapted to the target window position in the input image.

[0020] In a possible case, the window data is processed based on the first guide data, including:

[0021] Based on the first guide data, a correction value corresponding to the window data is determined, where the correction value indicates the correlation between different pixel points in the window data; and the window data is processed based on the correction value.

[0022] In this solution, the correlation between different pixel points of the window data is determined by considering the guide features of the spatial characteristics and the light sensitivity characteristics of the RAW image, so that the information in the window data that meets the light sensitivity characteristics and the spatial characteristics is extracted, the details and textures of the restored RAW image are improved, the artifacts are reduced, and the image quality is ensured.

[0023] Optionally, based on the first guide data, the correction value corresponding to the window data is determined, including: determining a target pixel point in the first guide data, where the target pixel point indicates the position of a pixel point obtained after processing the window data in the input image; and for each pixel point in the first guide data, the similarity between the pixel point and the target pixel point is calculated based on the data of the pixel point and the target pixel point in the first guide data, and the similarity is taken as the correction value of the pixel point.

[0024] In an example, the image restoration network includes an image encoding network and an image decoding network; and the m target layers are distributed in the image encoding network and / or the image decoding network.

[0025] In a possible implementation, the guide features of the image restoration network are determined by feature extraction based on the light sensitivity guide image and the spatial guide image, including: the light sensitivity guide image and the spatial guide image are subjected to feature extraction by a guide feature extraction module, and n guide feature maps output by the guide feature extraction module are determined.

[0026] In a possible implementation, the light sensitivity guide image corresponding to the RAW image is obtained, including: intermediate images corresponding to N pixel channels in the RAW image are determined; light sensitivity weights corresponding to the N pixel channels are determined based on the intermediate images corresponding to the N pixel channels; and the corresponding pixel channels are processed based on the light sensitivity weights corresponding to the N pixel channels to obtain the light sensitivity guide image.

[0027] In a possible implementation, determining the spatial guidance map corresponding to the RAW image comprises: obtaining, based on the RAW image, N position encoding maps respectively corresponding to the N pixel channels, the position encoding map indicating a position number of the corresponding pixel channel; determining N position weights respectively corresponding to the N pixel channels; and correcting the corresponding position encoding map based on the N position weights respectively corresponding to the N pixel channels to obtain the spatial guidance map.

[0028] In one example, the position encoding map further indicates an arrangement of the corresponding pixel channel in the target pixel of the color filter matrix.

[0029] In a possible implementation, the light sensing guidance map is an image obtained by shooting the target scene by the luminance sensor.

[0030] In this scheme, the light sensing guidance map is obtained by the luminance sensor, which is more consistent with the real scene, thereby ensuring the effect of image restoration.

[0031] In a second aspect, an embodiment of the present application provides an image restoration device, comprising:

[0032] An image acquisition module is configured to acquire a RAW image obtained by shooting a target scene by an image sensor, wherein the image sensor acquires the RAW image by using a front color filter matrix, the color filter matrix is arranged in a target pixel unit, the target pixel unit is arranged by N target pixels, target pixels at the same position of each target pixel unit in the color filter matrix form a pixel channel, and N pixel channels are formed in total;

[0033] A first guidance map determination module is configured to determine a light sensing guidance map corresponding to the RAW image, which indicates luminance information of the RAW image.

[0034] A second guidance map determination module is configured to determine a spatial guidance map corresponding to the RAW image, wherein the spatial guidance map indicates a position relationship of the N target pixels.

[0035] A feature extraction module is configured to perform feature extraction to determine guidance features of an image restoration network based on the light sensing guidance map and the spatial guidance map.

[0036] An image restoration module is configured to perform image restoration on the RAW image based on the guidance features and according to the image restoration network to obtain a restored image corresponding to the RAW image.

[0037] The beneficial effects of the present scheme are described above, and will not be repeated here.

[0038] In a possible implementation, the guidance feature indicates that the extraction of the photosensitive information and the information of the positional relationship of the N pixel channels in the RAW image can reflect the photosensitive difference of the N pixel channels and the difference of the pixel arrangement of the target pixels of the N pixel channels. The pixel arrangement indicates the distribution of the target pixels of the same pixel channel around the target pixel.

[0039] In a possible implementation, the image restoration network considers the photosensitive difference of the channels of the N target pixels of the RAW image and adopts different processing for the target pixels of the N pixel channels with different pixel arrangements to realize image restoration.

[0040] In a possible implementation, the image restoration module is configured to process at least part of the features extracted from the RAW image by the image restoration network based on the guidance feature.

[0041] The beneficial effects of the present solution are described above and will not be repeated here.

[0042] In a possible implementation, the number of target layers in the image restoration network is M, the guidance feature includes the guidance feature map corresponding to each of the m target layers in the M target layers, M is a positive integer greater than or equal to 2, and m is a positive integer less than or equal to M; the target layer is a convolution layer or a pooling layer; the image restoration module is configured to, during the process of image restoration of the RAW image by the image restoration network, process the window data of a target window of a first target layer on an input graph based on the guidance feature map corresponding to the first target layer when the first target layer is passed through; the input graph is a graph input into the first target layer, and the first target layer is any target layer in the m target layers.

[0043] The beneficial effects of the present solution are described above and will not be repeated here.

[0044] In an example, the image restoration module includes a first space determination unit, a second space determination unit, and a processing unit.

[0045] The first space determination unit is configured to determine the window data of the target window of the first target layer; the window data is the data of the position of the target window on the input graph.

[0046] The second space determination unit is configured to determine the first guidance data corresponding to the window data; the first guidance data is the data in the guidance feature map corresponding to the first target layer.

[0047] The processing unit is configured to process the window data based on the first guide data.

[0048] The beneficial effects of the present solution are described above and will not be repeated here.

[0049] In a possible case, the first guide data corresponding to the window data of the target window at different positions on the input image are different.

[0050] In a possible case, the guide feature map corresponding to the input image and the first target layer is of the same plane size, and the first guide data is data in the guide feature map corresponding to the first target layer, which is adapted to the position of the target window in the input image.

[0051] In a possible case, the processing unit is configured to determine a correction value corresponding to the window data based on the first guide data, wherein the correction value indicates the correlation between different pixel points in the window data, and the window data is processed based on the correction value.

[0052] The beneficial effects of the present solution are described above and will not be repeated here.

[0053] Optionally, the processing unit is configured to determine a target pixel point in the first guide data, wherein the target pixel point indicates the position of a pixel point obtained after processing the window data in the input image, and for each pixel point in the first guide data, the similarity between the pixel point and the target pixel point is calculated based on the data of the pixel point and the target pixel point in the first guide data, and the similarity is taken as the correction value of the pixel point.

[0054] In an example, the image recovery network includes an image encoding network and an image decoding network, and m target layers are distributed in the image encoding network and / or the image decoding network.

[0055] In an example, the feature extraction module is configured to perform feature extraction on the light sensing guide image and the space guide image by a guide feature extraction module, and determine n guide feature maps output by the guide feature extraction module.

[0056] In a possible implementation, the first guide image determination module includes a pixel recombination unit, a first weight determination unit, and a first guide image determination unit, wherein

[0057] The pixel recombination unit is configured to determine an intermediate image corresponding to each of the N pixel channels in the RAW image.

[0058] a first weight determination unit configured to determine a light-sensing weight corresponding to each of the N pixel channels based on the intermediate image corresponding to each of the N pixel channels;

[0059] a first guide map determination unit configured to obtain a light-sensing guide map by processing the corresponding pixel channel based on the light-sensing weight corresponding to each of the N pixel channels.

[0060] In a possible implementation, the second guide map determination module comprises a position encoding unit, a second weight determination unit and a second guide map determination unit, wherein,

[0061] The position encoding unit is configured to obtain a position encoding map corresponding to each of the N pixel channels based on the RAW image, the position encoding map indicating a position number of the corresponding pixel channel.

[0062] The second weight determination unit is configured to determine a position weight corresponding to each of the N pixel channels.

[0063] The second guide map determination unit is configured to obtain a spatial guide map by correcting the position encoding map corresponding to each of the N pixel channels based on the position weight corresponding to each of the N pixel channels.

[0064] In an example, the position encoding map further indicates an arrangement of the corresponding pixel channel at the target pixel of the color filter matrix.

[0065] In a possible implementation, the light-sensing guide map is an image obtained by a luminance sensor shooting the target scene.

[0066] In a third aspect, an embodiment of the present application provides an image recovery system, which can comprise an image sensor and an analysis device, wherein the analysis device is configured to execute the method provided in the first aspect.

[0067] In a fourth aspect, an embodiment of the present application provides an image recovery apparatus, which comprises at least one memory configured to store a program, and at least one processor configured to execute the program stored in the memory, and when the program stored in the memory is executed, the processor is configured to execute the method provided in the first aspect.

[0068] In a fifth aspect, an embodiment of the present application provides an image recovery apparatus, wherein the apparatus runs computer program instructions to execute the method provided in the first aspect. For example, the apparatus can be a chip or a processor.

[0069] In an example, the apparatus can comprise a processor, which can be coupled with a memory, read instructions in the memory and execute the method provided in the first aspect according to the instructions. The memory can be integrated in the chip or the processor, or can be independent of the chip or the processor.

[0070] In a sixth aspect, an embodiment of the present application provides a computer storage medium, which stores instructions, and the instructions, when executed on a computer, cause the computer to perform the method provided in the first aspect.

[0071] In a seventh aspect, an embodiment of the present application provides a computer program product containing instructions, and the instructions, when executed on a computer, cause the computer to perform the method provided in the first aspect. BRIEF DESCRIPTION OF DRAWINGS

[0072] Figure 1 FIG. 1 is a structural schematic diagram of an electronic device provided by an embodiment of the present application;

[0073] Figure 2 FIG. 2 is a structural schematic diagram of an image recovery device provided by an embodiment of the present application;

[0074] Figure 3a FIG. 3 is a schematic diagram of pixel arrangement of an R pixel channel provided by an embodiment of the present application;

[0075] Figure 3b FIG. 4 is a schematic diagram of pixel arrangement of a G pixel channel provided by an embodiment of the present application;

[0076] Figure 4 FIG. 5 is a flow schematic diagram of an image recovery method provided by an embodiment of the present application;

[0077] Figure 5a FIG. 6 is a flow schematic diagram of step 430 in FIG. 5; Figure 4

[0078] FIG. 7 is a flow schematic diagram of step 420 and step 430 in FIG. 5; Figure 5b Figure 4 FIG. 8 is an architecture schematic diagram of an image recovery model provided by an embodiment of the present application;

[0079] Figure 6a Figure 1 FIG. 9 is an architecture schematic diagram of an image recovery model provided by an embodiment of the present application;

[0080] Figure 6b FIG. 10 is a schematic diagram of pixel recombination shown in FIG. 8; Figure 2

[0081] FIG. 11 is a schematic diagram of pixel recombination shown in FIG. 9; Figure 7a Figure 6b Figure 1 FIG. 12 is a schematic diagram of pixel recombination shown in FIG. 10;

[0082] Figure 7b FIG. 13 is a schematic diagram of pixel recombination shown in FIG. 11; and Figure 6b Figure 2

[0083] Figure 7c FIG. 14 is a schematic diagram of pixel recombination shown in FIG. 12; and​​​​​​Figure 6a and Figure 6b schematic diagram of a guided convolution of a guided feature map;

[0084] Figure 7d is Figure 6a and Figure 6b schematic diagram of an image restoration network;

[0085] Figure 8 is Figure 4 flowchart of step 450 in FIG. 4;

[0086] Figure 9 is Figure 8 flowchart of step 451 in FIG. 4;

[0087] Figure 10 is a structural schematic diagram of an image restoration apparatus provided by an embodiment of the present application. DETAILED DESCRIPTION

[0088] In order to make the purpose, technical solutions and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments of the present application will be described below with reference to the drawings.

[0089] In the description of the embodiments of the present application, the words “exemplary”, “for example”, or “for instance” are used to mean serving as an example, instance or illustration. Any embodiment or design solution described as “exemplary”, “for example” or “for instance” in the embodiments of the present application should not be interpreted as being more advantageous or superior than other embodiments or design solutions. Rather, the use of the words “exemplary”, “for example” or “for instance” is intended to present the relevant concept in a specific manner.

[0090] In the description of the embodiments of the present application, the term “and / or” merely describes an association relationship of associated objects, and can represent three relationships, for example, A and / or B, which can represent three cases of existence of A alone, existence of B alone, and existence of A and B simultaneously. In addition, unless otherwise specified, the term “multiple” means two or more. For example, multiple systems mean two or more systems, and multiple terminals mean two or more terminals.

[0091] In addition, the terms “first” and “second” are used for description purposes only, and should not be interpreted or implied to indicate or suggest relative importance or implicitly indicate the indicated technical features. Therefore, the features defined with “first” and “second” can explicitly or implicitly include one or more features. The terms “include”, “contain”, “have” and their variants mean “include but are not limited to”, unless otherwise specifically emphasized.

[0092] The image recovery method provided by the embodiment of the present application can be applied to an electronic device, which can be various terminals such as a smart phone, a desktop computer, a notebook computer, a palm computer, a tablet computer, smart glasses, a wearable device, a camera, a video camera, and the like. In addition, the electronic device provided by the embodiment of the present application can also provide cloud services, and can be a physical server in the cloud.

[0093] As shown in Figure 1 , it is an architecture schematic diagram of an exemplary electronic device 100 provided by the embodiment of the present application. Figure 1 The structure schematic diagram of the electronic device embodiment provided by the present application is shown in Figure 1 , the electronic device 100 in the embodiment includes a processor 111, a memory 112, a nonvolatile memory 1121, a random access memory 1122, an antenna 1, an antenna 2, a communication module 113, an audio module 114, a speaker 114A, a receiver 114B, a microphone 114C, an earphone interface 114D, a display screen 115, a camera 116, a sensor module 117, an input device 118, and a power supply 119.

[0094] It can be understood that the structure shown in the embodiment of the present application does not constitute a specific limitation on the electronic device 100. In other embodiments of the present application, the electronic device 100 can include more or fewer components than shown, or combine certain components, or split certain components, or different arrangement of components. The components shown can be implemented in hardware, software, or a combination of software and hardware.

[0095] The specific introduction of certain components of the electronic device 100 will be described below. Figure 1

[0096] ​The processor 111 can include one or more processors, for example, the processor 111 can include one or more of an application processor (AP), a modem, a graphics processing unit (GPU), an image signal processor (ISP), a controller, a video codec, a digital signal processor (DSP), a baseband processor, and / or a neural-network processing unit (NPU), etc. Among them, different processors can be independent devices, or can be integrated in one or more processors. For example, the processor 111 can process the content required to be displayed on the display window of the application program on the electronic device 100. For example, the controller can generate operation control signals according to instruction operation codes and timing signals, complete instructions and control the execution of instructions.

[0097] In one example, a memory can also be provided in the processor 111 for storing instructions and data. In some examples, the memory in the processor 111 is a cache memory. The memory can hold instructions or data that the processor 111 has just used or is using in a loop. If the processor 111 needs to use the instructions or data again, it can be directly called from the memory, avoiding repeated access, reducing the waiting time of the processor 111, and improving the efficiency of the system.

[0098] The memory 112 can include an internal memory for storing computer executable program codes including instructions. The processor 111 performs various functional applications and data processing of the electronic device 100 by executing the instructions stored in the internal memory. The internal memory can include a program storage area and a data storage area. The program storage area can store an operating system, at least one application program (e.g., a sound play function, an image play function, etc.) required for at least one function, a client, etc. The data storage area can store data (e.g., an audio signal, a phone book, etc.) created during the use of the electronic device 100, etc. In addition, the internal memory can include a non-volatile memory 1121 such as at least one of a disk storage device, a flash memory device, a universal flash storage (UFS), etc., and a random access memory 1122 commonly referred to as a memory such as a static random access memory (SRAM), a dynamic random access memory (DRAM), a synchronous dynamic random access memory (SDRAM), a double data rate synchronous dynamic random access memory (DDR SDRAM), an enhanced SDRAM (ESDRAM), a synchlink DRAM (SLDRAM), and a direct rambus RAM (DR RAM).

[0099] In one example, the memory 112 further includes an external memory card such as a Micro SD card connected through an external memory interface to implement an extension of the storage capacity of the electronic device 100. The external memory card communicates with the processor 111 through the external memory interface to implement a data storage function. For example, files such as music, video, etc. are saved in the external memory card.

[0100] The wireless communication function of the electronic device 100 can be implemented through the antenna 1, the antenna 2, the communication module 113, a modem, and a baseband processor, etc. The communication module includes a wireless communication module and a wired communication module.

[0101] The antenna 1 and the antenna 2 are used to transmit and receive electromagnetic wave signals. Each antenna in the electronic device 100 can be used to cover a single or multiple communication bands. Different antennas can also be multiplexed to improve the utilization of the antennas. For example, the antenna 1 can be multiplexed as a diversity antenna for a wireless local area network. In some other examples, the antennas can be used in combination with a tuning switch.

[0102] The mobile communication module can provide a solution for wireless communication including 2G / 3G / 4G / 5G, etc. applied to the electronic device 100. The mobile communication module can include at least one filter, a switch, a power amplifier, a low noise amplifier (LNA), etc. The mobile communication module can receive an electromagnetic wave by at least two antennas including the antenna 1, and perform filtering, amplification, etc. on the received electromagnetic wave, and transfer the processed electromagnetic wave to a modem to be demodulated. The mobile communication module can also amplify a signal modulated by the modem, and radiate the amplified signal as an electromagnetic wave through the antenna 1. In some examples, at least part of the functional modules of the mobile communication module can be disposed in the processor 111. In some examples, at least part of the functional modules of the mobile communication module can be disposed in the same device as at least part of the modules of the processor 111.

[0103] The wireless communication module can provide a solution for wireless communication including wireless local area networks (WLAN) (e.g., wireless fidelity (Wi-Fi) network), Bluetooth (BT), global navigation satellite system (GNSS), frequency modulation (FM), near field communication (NFC), infrared (IR) technology, etc. applied to the electronic device 100. The wireless communication module can be one or more devices integrating at least one communication processing module. The wireless communication module receives an electromagnetic wave through the antenna 2, performs frequency modulation and filtering on the electromagnetic wave signal, and transmits the processed signal to the processor 111. The wireless communication module can also receive a signal to be transmitted from the processor 111, perform frequency modulation and amplification on the signal, and radiate the processed signal as an electromagnetic wave through the antenna 2.

[0104] The electronic device 100 can implement an audio function through the audio module 114, the speaker 114A, the receiver 114B, the microphone 114C, the earphone interface 114D, the application processor, etc. For example, music playback, recording, etc.

[0105] The audio module 114 is used to convert digital audio information into an analog audio signal output, and is also used to convert an analog audio input into a digital audio signal. The audio module 114 can also be used to encode and decode an audio signal. In some examples, the audio module 114 can be disposed in the processor 111, or part of the functional modules of the audio module 114 can be disposed in the processor 111.

[0106] The speaker 114A, also called a "loudspeaker", is used to convert an audio electrical signal into a sound signal. The electronic device 100 can listen to music or listen to a hands-free call through the speaker 114A.

[0107] The receiver 114B, also called a "earpiece", is used to convert an audio electrical signal into a sound signal. When the electronic device 100 answers a call or a voice message, the receiver 114B can be held close to a human ear to listen to the voice.

[0108] The microphone 114C, also called a "microphone", "sound pickup", is used to convert a sound signal into an electrical signal. When making a call or sending a voice message, a user can speak into the microphone 114C through the human mouth to input a sound signal into the microphone 114C. The electronic device 100 can be provided with at least one microphone 114C. In other examples, the electronic device 100 can be provided with two microphones 114C, in addition to collecting sound signals, noise reduction functions can also be realized. In other embodiments, the electronic device 100 can also be provided with three, four or more microphones 114C, in addition to collecting sound signals, noise reduction, and can also identify the source of the sound, realize the function of directional recording, etc.

[0109] The earphone interface 114D is used to connect a wired earphone. The earphone interface 114D can be a USB interface 130, or a 3.5mm open mobile terminal platform (OMTP) standard interface, a cellular telecommunications industry association of the USA (CTIA) standard interface.

[0110] The electronic device 100 realizes the display function through the GPU, the display screen 115, the memory 112, the digital-to-analog converter, and the application processor, etc. The GPU is a microprocessor for image processing, connected to the display screen 115 and the application processor. The GPU is used to perform mathematical and geometric calculations for graphics rendering. The processor 111 can include one or more GPUs that execute program instructions to generate or change display information. The memory 112 includes a video memory that is used to store data processed by the GPU or data processed by the application processor, which represents information for each pixel that can be output to the display screen 115; the video memory can be independent and does not need to occupy the random access memory 1122; it can also share part of the random access memory 1122 as a video memory; it can also be an independent video memory and a shared part of the random access memory 1122. The digital-to-analog converter is used to convert the video memory data into an analog signal, and the display screen 115 is used to realize the display based on the analog information.

[0111] Specifically, the display screen 115 is configured to display images, videos, and the like. The display screen 115 includes a display panel. The display panel can be a liquid crystal display (LCD), an organic light-emitting diode (OLED), an active-matrix organic light-emitting diode (AMOLED), a flex light-emitting diode (FLED), a Miniled, a MicroLed, a Micro-oLed, a quantum dot light emitting diode (QLED), or the like. In some examples, the electronic device 100 can include one or more display screens 115. In one example, the display screen 115 can be configured to display an interface of an application, display a display window of an application, and the like. Notably, here, the electronic device 100 has the display screen 115. Alternatively, the display screen 115 can be directly operated, and specifically, the electronic device 100 can be a mobile phone, a tablet, or the like. Alternatively, the electronic device 100 is connected with an external operating device such as a keyboard 117 and a mouse 118, data is input through the keyboard 117, and the content displayed on the display screen 115 is operated through the mouse 118. Specifically, the electronic device 100 can be a notebook computer, a desktop computer, or the like. In addition, the exchange of data and control information between the display screen 115 and the processor 111 is implemented through a display controller in the I / O subsystem.

[0112] The camera 116 is configured to capture images or videos, and can be triggered to start by an application instruction to implement a photographing or video shooting function, such as capturing pictures or videos of any scene. The camera can include an imaging lens, a filter, an image sensor, and the like. Light emitted or reflected by an object enters the imaging lens, passes through the filter, and is finally focused on the image sensor. The imaging lens is mainly used to focus the light emitted or reflected by all objects (which can also be referred to as a scene to be photographed, a target scene, or a scene image that a user expects to capture) in a photographing angle of view; the filter is mainly used to filter out unnecessary light waves (such as light waves other than visible light, such as infrared) in the light; and the image sensor is mainly used to perform photoelectric conversion on the received light signals, convert them into electrical signals, and input them to the processor 111 for subsequent processing. The camera can be located on the front of the electronic device 100 or on the back of the electronic device 100, and the specific number and arrangement of the camera can be flexibly determined according to the needs of designer or manufacturer strategies, which are not limited in the present application.

[0113] The sensor module 117 can include a pressure sensor, a gyro sensor, a barometric pressure sensor, a magnetic sensor, an acceleration sensor, a distance sensor, a proximity light sensor, a fingerprint sensor, a temperature sensor, a touch sensor, an ambient light sensor, a bone conduction sensor, etc. Here, the touch sensor is also referred to as a "touch panel". The touch sensor can be disposed on the display screen 115, and the touch sensor and the display screen 115 can constitute a touch screen, also referred to as a "touch screen". The touch sensor is configured to detect touch operation data acting on or near the touch sensor. The touch sensor can transmit the detected touch operation data to the application processor to determine a touch event type. Visual output related to the touch operation data can be provided through the display screen 115. The sensor 180 can include an image sensor, a motion sensor, a proximity sensor, an ambient noise sensor, a sound sensor, an accelerometer, a temperature sensor, a gyroscope, or other types of sensors, and various combinations thereof. The processor 111 drives the sensor module to receive various information such as audio signals, image signals, motion information, etc. through a sensor controller in the I / O subsystem, and the sensor module transmits the received information to the processor 111 for processing.

[0114] The input device 118 can include an external input device such as a key and a mouse. The key includes a power-on key, a volume key, an input keyboard, etc. The key can be a mechanical key. It can also be a touch key. The electronic device 100 can receive key input, generate key signal input related to user settings and function control of the electronic device 100. In addition, data and control information exchange between the input device 118 and the processor 111 is achieved through an input device controller in the I / O subsystem.

[0115] Further, the electronic device 100 can further include a power supply 119 to supply power to other components of the electronic device 100 including 111-118, which can be a rechargeable or non-rechargeable lithium-ion battery or a nickel-hydrogen battery. Further, when the power supply is a rechargeable battery, the power management system can be coupled with the processor 111, so that the power management system can achieve functions such as management of charging, discharging, and power consumption adjustment.

[0116] Of course, in order to simplify, Figure 1 Only some of the components related to the embodiments of the present application in the electronic device 100 are shown in the figure, and components such as buses, input / output interfaces, etc. are omitted. In addition, according to specific application circumstances, the electronic device 100 can also include any other appropriate components. Those skilled in the art can understand that, Figure 1 The electronic device 100 is only an example, and does not constitute a limitation on the electronic device, which can include more or fewer components than shown, or combine certain components, or different components.

[0117] AsFigure 2 The diagram shown is a hardware architecture diagram of an exemplary image restoration device provided in an embodiment of this application. The image restoration device 200 can, for example, be a processor chip. Figure 2 The hardware architecture diagram shown may be Figure 1 An exemplary architecture diagram of the processor 130 is shown in the present application embodiment, and the image processing method provided in this application embodiment can be applied to the processor chip.

[0118] refer to Figure 2 The device 200 includes: at least one central processing unit 211, a memory 212, a graphics processor 213, a neural network processor 214, a memory bus 215, a receiving interface 216, and a transmitting interface 217, etc. Although Figure 2 As not shown, the device 200 may also include an application processor (AP), a decoder, and a dedicated video or image processor and a microcontroller unit (MCU).

[0119] The various parts of the device 200 are coupled together by connectors. For example, the connectors include various interfaces, transmission lines or buses, etc. These interfaces are usually electrical communication interfaces, but may also be mechanical interfaces or other forms of interfaces. This embodiment does not limit them.

[0120] Optionally, the central processing unit 211 can be a single-core or multi-core processor; alternatively, the central processing unit 211 can be a processor group consisting of multiple processors, which are coupled to each other through one or more buses. The receiving interface can be a data input interface; in one optional case, the receiving and transmitting interfaces can be High Definition Multimedia Interface (HDMI), V-By-One interface, Embedded Display Port (eDP), Mobile Industry Processor Interface (MIPI), or DisplayPort (DP), etc. The memory can be referred to the foregoing description of the memory 112 section.

[0121] In an optional case, the above-mentioned parts are integrated on the same chip; in another optional case, the central processor 211, the graphic processor 213, the decoder, the receiving interface 216 and the sending interface 217 are integrated on a chip, and the parts inside the chip access the external memory through a bus. The special video / graphic processor can be integrated with the central processor 211 on the same chip, or can exist as a separate processor chip, for example, the special video / graphic processor can be a special ISP. In an optional case, the neural network processor 214 can also exist as an independent processor chip. The neural network processor 214 is used to implement various neural network or deep learning related operations. Optionally, the image processing method provided in the embodiments of the present application can be implemented by the graphic processor 213 or the neural network processor 214, or can be implemented by a special graphic processor.

[0122] The chip involved in the embodiments of the present application is a system manufactured on the same semiconductor substrate by integrated circuit technology, also called a semiconductor chip, which can be a collection of integrated circuits formed on a substrate (usually a semiconductor material such as silicon) by integrated circuit technology, and the outer layer is usually packaged by a semiconductor packaging material. The integrated circuit can include various functional devices, each of which includes a logic gate circuit, a metal oxide semiconductor (MOS) transistor, a bipolar transistor or a diode transistor, and can also include a capacitor, a resistor or an inductor and other components. Each functional device can work independently or under the action of necessary driving software, and can realize various functions such as communication, calculation or storage.

[0123] Before introducing the technical solutions provided in the present application, the RAW image is first introduced.

[0124] The RAW image is an unprocessed original image obtained by an image sensor, and each pixel of the RAW image represents the intensity of only one color. The image sensor for generating the RAW image includes a light sensing element (CMOS or CCD described above) and a color filter matrix, and the color filter matrix is overlaid on the light sensing element. The color filter matrix is arranged by a target pixel matrix, and the number and color of target pixels in the target pixel matrix can be determined according to actual needs. It is worth noting that the target pixel matrix is the smallest repeating unit.

[0125] The image sensor can be a CMOS sensor or a CCD sensor, for example. The color format of the RAW image is determined by a color filter array (CFA) placed in front of the image sensor (i.e., the filter mentioned above), and the RAW image can be an image obtained in various CFA formats. For example, the RAW image can be a Bayer image in an RGGB format, in which each cell represents a target pixel, R represents a red pixel, G represents a green pixel, and B represents a blue pixel. The smallest repeating unit of the Bayer image is a 2x2 array, i.e., a target pixel matrix, and the 2x2 array unit includes four target pixels of R, G, G, and B. Alternatively, the RAW image can be an image in a red yellow yellow blue (RYYB) format or an image in an XYZW format, in which XYZW represents an image format including four components, and X, Y, Z, and W each represent a component, such as a Bayer image in a Red Green Blue Infrared (RGBIR) arrangement or a Bayer image in a Red Green Blue White (RGBW) arrangement. The RAW image can also be an image in a Quad arrangement, in which four target pixels of the same color are arranged together, and it is noted that the Quad arrangement is an RGGB arrangement if the four target pixels of each color are considered as a whole.

[0126] The RAW image has spatial characteristics and light sensing characteristics. Since the spatial characteristics and the light sensing characteristics are described for a pixel channel of the RAW image, the pixel image of the RAW image is described first. The pixel channel indicates a channel formed by target pixels of the same position in a plurality of target pixel matrices, i.e., the smallest repeating unit, in a color filter array.

[0127] For example, assume that the RAW image is an image captured by an RGGB sensor. Since the color filter array of the RGGB sensor is arranged in an RGGB pixel matrix, the RAW image has four pixel channels of R, G, G, and B. Here, R represents a red color, G represents a green color, and B represents a blue color.

[0128] For example, assume that the RAW image is an image captured by an RGBW sensor. Since the color filter array of the RGBW sensor is arranged in an RGBW pixel matrix, the RAW image has four pixel channels of R, G, B, and W. Here, W represents a white color.

[0129] Assume that the RAW image is an image captured by an RYYB sensor. Since the color filter matrix in the RYYB sensor is arranged in an RYYB pixel matrix, the RAW image has four pixel channels of R, Y, Y, and B. Here, Y represents yellow.

[0130] For spatial features, the spatial distribution characteristics of different pixel channels of the RAW image of the image sensor are different, and due to the spatial distribution characteristics, the processing operations of pixels at different positions should be different. As shown in FIGS. 1A and 1B, taking the most common Bayer format data as an example, Figure 3a and Figure 3b As shown in FIGS. 1A and 1B, taking the most common Bayer format data as an example, Figure 3a the image of the R channel (black in the figure represents red), there are four position cases of interpolation operations, and the data distribution around different positions (the gray dashed box) is different, and different processing operations should be adopted; Figure 3b the image of the G channel (black in the figure represents green), there are two position cases of interpolation operations, and the data distribution around different positions (the gray dashed box) is different, and different processing operations should be adopted. In addition, with the development of image sensors, image sensors have more complex spatial distribution characteristics, and how to more effectively utilize the spatial characteristics of the RAW image becomes particularly important.

[0131] For light sensing features, different pixel channels of the image sensor receive different wavebands of visible light, resulting in different signal-to-noise ratios, brightness, and semantic content between different channels; further, for different image sensors, the imaging modes of different image sensors are different, resulting in large differences in the light sensing characteristics of RAW of different image sensors. For example, assume that the W channel of the RGBW sensor has the highest peak signal-to-noise ratio, then better recovery effect can be obtained through the data of the W channel; but in the RAW image of the sub-pixel sensor, due to overexposure, directly selecting the channel with the highest peak signal-to-noise ratio as the guide feature may cause loss of part of the image content, making the recovery effect worse; therefore, the light sensing characteristics cannot be simply judged according to experience, and the characteristics of different image sensors need to be considered.

[0132] In the field of computer vision, photographing has become one of the most commonly used functions of various mobile terminal devices such as mobile phones, tablet computers, smart glasses, wearable devices, etc. At present, the image obtained through an image sensor (such as a camera) is an unprocessed RAW image, and then a series of image processing operations are performed on the RAW image to convert it into a standard image for display. Among them, this series of image processing operations need to restore as much detail and clarity of the real content as possible, and the higher the image quality of the restored image, the closer it is to the real shooting content. It should be understood that the format of the standard image finally displayed on the display device can also be a Red Green Blue (RGB) format, or other image formats such as YUV color images, YCbCr color images, or grayscale images, etc. The embodiments of the present application take the image finally displayed on the display device as an example to illustrate the RGB image.

[0133] However, due to the performance and power consumption of the image sensor, the image quality of the RAW image collected by the image sensor is still not high enough, and there are problems such as low resolution, serious noise, insufficient brightness, missing details, blur, etc. Therefore, the quality recovery of the RAW image is a very urgent need, which can greatly improve the visual effect and help improve the accuracy of later visual tasks such as target detection, scene segmentation, etc.

[0134] In recent years, deep learning methods, especially CNN (Convolutional Neural Networks) based convolutional neural network methods, have achieved the best performance in the field of image restoration, gradually surpassing traditional algorithms. However, due to the huge differences in RAW images collected by different manufacturers / models of image sensors, these differences will directly affect the robustness and generalization of RAW image restoration algorithms.

[0135] In summary, it is urgent to propose more effective RAW image restoration methods for multiple sensor types.

[0136] At present, image restoration can be achieved through the following three implementation methods.

[0137] Implementation method 1: first perform Pixel Shuffle (pixel reorganization) processing on the input RAW image, and obtain the restored image through the CNN network.

[0138] However, the above method has insufficient details in the restoration of the RAW image, serious artificial traces, and poor restoration effect.

[0139] Implementation method 2: first restore the channel with high sensitivity and high signal-to-noise ratio to obtain a restored partial channel image; then, further fuse and process the RAW image according to the restored channel image to finally obtain a restored image.

[0140] The above method is prone to cause detail loss and serious artifacts due to only using the photosensitive characteristics.

[0141] Implementation 3: The image and the extracted guide information are recovered simultaneously through the CNN network, the guide image is recovered based on the extracted guide information, and the final recovery result is obtained by inputting the RAW image adaptive recovery generated.

[0142] The guide information extracted by the above method needs to be recovered into a guide image, which limits the guide effect.

[0143] In order to solve the above problems of image recovery, based on this, an image recovery method is provided in the embodiments of the present application. Here, only the method is briefly described, and the detailed content of the method is described below.

[0144] The method considers the spatial features and photosensitive characteristics of the RAW image to jointly guide the image recovery network to realize the recovery of the RAW image, and improves the image quality after the RAW image recovery. The image recovery network can comprehensively consider the photosensitive difference of different pixel channels of the RAW image and the pixel distribution difference of the target pixels under the guidance of the spatial features and the photosensitive characteristics. The pixel distribution indicates the distribution of the target pixels belonging to the same pixel channel around the target pixel.

[0145] More specifically, the embodiments of the present application extract the guide features of the spatial features and the photosensitive characteristics of the RAW image, and correct the features extracted in the image recovery network based on the guide features, so as to realize the image recovery.

[0146] It should be understood that image recovery refers to a technology of recovering lost parts of an image and reconstructing them based on image information. Through image recovery, it can be tried to estimate the original image information, repair and improve the damaged areas, so as to improve the visual quality of the image.

[0147] Next, an image recovery method provided by the embodiments of the present application will be described in detail.

[0148] Figure 4 is a flowchart of the image recovery method provided by the embodiments of the present application. The embodiments can be applied on an electronic device, and specifically can be applied on a server or a general computer.

[0149] As shown in Figure 4 The image recovery method provided by the embodiments of the present application at least includes the following steps:

[0150] Step 410, obtaining a RAW image obtained by an image sensor for shooting a target scene, the RAW image including N pixel channels.

[0151] It should be understood that if the execution subject of the image restoration method is a processor chip as shown in Figure 2 , the RAW image can be obtained through the receiving interface 216, which is obtained by the camera 116 of the electronic device 100; if the execution subject of the image processing is the electronic device 100 as shown in Figure 1 , the RAW image can be obtained through the camera 116.

[0152] Step 420, determine the light sensing guide map corresponding to the RAW image; wherein the light sensing guide map indicates the brightness information of the RAW image.

[0153] According to a possible implementation, based on the fact that the brightness sensor and the image sensor shoot the target scene at the same time, the light sensing guide map G L corresponding to the RAW image collected by the image sensor is obtained.

[0154] According to a possible implementation, the light sensing of the RAW image in different pixel channels can be analyzed, and the signal-to-noise ratio, brightness, semantic content and other information of different pixel channels are comprehensively considered, the better-performing pixel channel is found, the information of the better-performing pixel channel is focused on, and the light sensing guide map G L is obtained. Here, the light sensing guide map G L can guide the image restoration network to strengthen the light information of the better-performing pixel channel and weaken the light information of the worse-performing pixel channel. For example, the weights of the N pixel channels of the RAW image can be determined, the N pixel channels are processed (corrected or fused) based on the weights of the N pixel channels, and the light sensing guide map G L is obtained.

[0155] Step 430, determine the spatial guide map corresponding to the RAW image; wherein the spatial guide map indicates the positional relationship of the N pixel channels.

[0156] In the embodiments of the present application, the pixel distribution of each pixel channel of the RAW image is analyzed, the data distribution of different pixel channels is comprehensively considered, the pixel channel that is more important for image restoration is found, the data distribution of the better-performing pixel channel is focused on, and the spatial guide map G P is obtained. Here, the spatial guide map G P can guide the image restoration network to strengthen the data distribution of the better-performing pixel channel and weaken the data distribution of the worse-performing pixel channel.

[0157] For example, the N pixel channels are respectively positionally encoded to obtain N positionally encoded values of the N pixel channels, such as 1, 2, 3, …, N. Optionally, for each of the N pixel channels, the pixel channel is arranged according to a pixel pattern of the pixel channel, and the pixel value of the pixel channel is replaced by the corresponding positionally encoded value to obtain a positionally encoded image. Subsequently, the N pixel channels of the RAW image can be determined to have respective position weights, and the N positionally encoded images are processed (corrected or fused) based on the respective position weights of the N pixel channels to obtain a spatial guidance image G P . Optionally, the N pixel channels of the RAW image are replaced by the corresponding positionally encoded values to obtain a spatial guidance image G P In some possible cases, the N pixel channels of the RAW image can be determined to have respective position weights, and the positionally encoded values in the spatial guidance image G P are corrected based on the respective position weights of the N pixel channels.

[0158] Step 440: Based on the light sensing guidance image and the spatial guidance image, feature extraction is performed to determine a guidance feature of the image restoration network.

[0159] In the embodiments of the present application, the data distribution and light sensing of the different pixel channels of the RAW image are comprehensively analyzed, and the light sensing feature and the spatial feature are extracted to obtain the guidance feature.

[0160] In one possible implementation, the guidance feature indicates the extracted light sensing information, and the feature obtained from the position relationship information of the N pixel channels of the RAW image can reflect the light sensing difference of the N pixel channels and the difference in pixel arrangement of the target pixels of the N pixel channels. The pixel arrangement indicates the distribution of the target pixels belonging to the same pixel channel around the target pixel. For example, the pixel arrangement can be understood as the distribution of the pixels around 1, 2, 3, and 4 in Figure 3a and Figure 3b .

[0161] Step 450: Based on the guidance feature, the RAW image is restored by the image restoration network to obtain a restored image corresponding to the RAW image.

[0162] In one possible implementation, the image restoration network considers the light sensing difference of the N target pixels of the channels of the RAW image, and different processing is adopted for the target pixels of the N pixel channels having different pixel arrangements to realize image restoration. The pixel arrangement indicates the distribution of the target pixels belonging to the same pixel channel around the target pixel. For example, the pixel arrangement can be understood as the distribution of the pixels around 1, 2, 3, and 4 in Figure 3a and Figure 3b .

[0163] In some possible cases, based on guided features, at least some of the features extracted by the image restoration network can be modified so that the modified features have photosensitive and spatial properties, thereby better restoring the details and texture of the image and reducing the generation of artifacts.

[0164] In summary, the embodiments of this application combine the spatial and photosensitive characteristics of RAW images to assist the image restoration network in restoring RAW images, improving the details and textures of the restored RAW images, reducing artifacts, ensuring image quality, and thus improving visual effects. In addition, it helps to improve the accuracy of subsequent visual tasks, such as object detection and scene segmentation.

[0165] Figure 5a As shown Figure 4 The flowchart of step 430 in the illustrated embodiment is shown.

[0166] like Figure 5a As shown above, in the above Figure 4 Based on the illustrated embodiment, in this embodiment of the application, step 430 may specifically include the following steps:

[0167] Step 431: Based on the RAW image, obtain the position encoding map corresponding to each of the N pixel channels. The position encoding map indicates the position number value of the corresponding pixel channel.

[0168] Step 432: Determine the position weights corresponding to each of the N pixel channels.

[0169] Step 433: Correct the corresponding positional encoding map based on the positional weights of each of the N pixel channels to obtain the spatial guidance map.

[0170] In this embodiment of the application, by considering the positional weights of different pixel channels, the spatial guiding map G can be ensured. P The reference value. Subsequently, based on the spatial guidance graph G P The extracted guiding features can give more attention to pixel channels with higher positional weights.

[0171] In one possible scenario, the photoguided image G L The same size as the RAW image. For example, such as... Figure 6a As shown, the photoguide map G is determined mainly through two implementation methods. L Implementation method A1: Based on the simultaneous capture of the target scene by the brightness sensor and the image sensor, the photosensitive guide map G corresponding to the RAW image acquired by the image sensor is obtained. L Implementation method A2: Obtain the photoguided image G by processing the RAW image. L .

[0172] Correspondingly, in step 431, as shown in Figure 6a the spatial guidance map G P .

[0173] Implementation B1, each of the N pixel channels is position encoded to obtain a position encoding value of each of the N pixel channels, such as 1, 2, 3, …, N; the pixels of the N pixel channels in the RAW image are replaced by the corresponding position encoding values to obtain the spatial guidance map G P .

[0174] Implementation B2, each of the N pixel channels is position encoded to obtain a position encoding value of each of the N pixel channels, such as 1, 2, 3, …, N; the photosensitive guidance map G L or the RAW image is processed to obtain a weight value of each of the N pixel channels; for each pixel channel of the N pixel channels, the position encoding value is corrected based on the weight value of the pixel channel to obtain a corrected position encoding value; the pixels of the N pixel channels in the RAW image are replaced by the corresponding corrected position encoding values to obtain the spatial guidance map G P . Wherein, the processing of the photosensitive guidance map G L or the RAW image is achieved by a deep learning network, and the structure of the deep learning network is not specifically limited in the embodiments of the present application, and can be determined according to actual needs.

[0175] In one possible case, the photosensitive guidance map G L and the RAW image have different sizes. For example, as shown in Figure 6b , the RAW image is pixel reorganized to obtain N intermediate images corresponding to the N pixel channels respectively, the weight values corresponding to the N intermediate images are determined, and then the corresponding intermediate channels are corrected based on the weight values corresponding to the N intermediate images to obtain the photosensitive guidance map G L .

[0176] Figure 5b The flowchart of step 420 in the embodiment shown in Figure 4 is shown.

[0177] As shown in Figure 5b , based on the above-described Figure 4 embodiment, in the embodiments of the present application, step 420 can specifically include the following steps:

[0178] Step 421, determining intermediate images corresponding to the N pixel channels in the RAW image respectively.

[0179] The RAW image is pixel reorganized to obtain intermediate images of the N pixel channels respectively

[0180] For convenience of description and distinction, the length and width of the input RAW image are denoted as H and W respectively.

[0181] Here, the RAW image (RGGB image, assuming the dimension of H*W*1) is subjected to a pixel reorganization operation to obtain the input feature F0, including N intermediate images of respective pixel channels.

[0182] For example, when the input RAW image (assuming the dimension of H*W*1) is an RGGB image, the RAW image is subjected to a pixel reorganization operation to obtain the input feature F0 (dimension of (H / 2)*(W / 2)*4), where N=4, and the input feature F0 is composed of four intermediate images, each of which represents an image of a pixel channel, and the four pixel channels can be R, G, B and W channels. It should be noted that the minimum repeating unit of the RGGB format Bayer image includes four pixels of R, G, G and B, and the four pixels R, G, G and B in each minimum repeating unit in the RAW image are split and rearranged to obtain four different intermediate images, i.e., a RAW image of W*H is split into four intermediate images of H / 2*W / 2. That is, when the input RAW image is an RGGB format Bayer image, four intermediate images of H / 2*W / 2 can be obtained after pixel reorganization, each of which contains only one color component. Specifically, the four intermediate images are an R image belonging to a first channel, a G image belonging to a second channel, a G image belonging to a third channel and a B image belonging to a fourth channel. As shown in FIG. 2, the size of the RGGB image is 6*6*1, and for the pixel R in the minimum repeating unit, the pixel R in each minimum repeating unit in the RGGB image is extracted and arranged to obtain the intermediate image of the pixel R, and the pixels G, G and B are similar and will not be described again. Figure 7a

[0183] For example, when the input RAW image is an RYYB image or an XYZW image, the input feature F0 (dimension of (H / 2)*(W / 2)*4) is obtained after pixel reorganization, where N=4, and the difference from the RGGB image is only in the pixel channels.

[0184] ​For example, when the input RAW image is a quad-arranged image, after pixel recombination, the input feature F0 (dimension (H / 4)*(W / 4)*16) is obtained, where N=16. It should be noted that the smallest repeating unit of a quad-arranged image includes 4 R, G, G, and B pixels, totaling 16 pixels. After pixel recombination of a W*H quad-arranged RAW image, 16 intermediate images of size H / 4×W / 4 are obtained, where each intermediate image belongs to one channel. That is, when the input RAW image is a quad-arranged image containing 16 pixels, pixel recombination yields intermediate images belonging to 16 channels. Figure 7b As shown, the RAW image is arranged in a Quad pattern with a size of 12*12*1. The smallest repeating unit contains four pixels R, referred to as R1, R2, R3, and R4 for ease of description and distinction. For R1, pixels R1 in each smallest repeating unit of the RAW image are extracted and arranged to obtain the intermediate image of pixel R1. Pixels R2, R3, and R4 are similar and will not be described further. Each of the smallest repeating units contains four pixels G, G, and B. Following the above method, the corresponding intermediate images can be obtained. In an optional scheme, the number of R, G, G, and B pixels in the smallest repeating unit of the Quad-arranged image can also be 6, 8, or other numbers. Correspondingly, after pixel recombination, intermediate images belonging to 24 channels or 32 channels can be obtained.

[0185] In summary, it should be understood that the number of intermediate images is equal to the number of target pixels contained in the smallest repeating unit of the RAW image.

[0186] Step 422: Extract information from the intermediate images of each of the N pixel channels using the first weight extraction module to determine the photosensitive weights of each of the N pixel channels.

[0187] According to a feasible implementation, the first weight extraction module is used to extract information by pooling the above-mentioned input features F0, and obtain the weights ω of N pixel channels. L (Dimension 1×1×N), where each value represents the weight of a pixel channel.

[0188] In one example, the first weight extraction module uses different pooling methods to pool the input feature F0, obtaining the pooled features F under different pooling methods. p (Dimension 1*1*N), pooling features F after different pooling methods p The features are fused along the channel dimension (1*1*N) to obtain fused feature F1 (1*1*N). Subsequently, fused feature F1 is passed through a fully connected layer to obtain the weights ω of N pixel channels. L(Dimension 1*1*N). Here, there can be multiple different pooling methods, such as two: one is max pooling, and the other is global average pooling.

[0189] For example, such as Figure 6b As shown, when the RAW image is an RGGB image, the first weight extraction module extracts information by pooling the input feature F0 (with dimensions of (H / 2)*(W / 2)*4) to obtain the weights ω of the four channels. L (Dimension 1×1×4), each value represents the weight of a pixel channel, yielding the weight values ​​for the R, G, B, and W channels. Subsequently, the input feature F0 (dimension (H / 2)*(W / 2)*4) is pooled using max pooling and global average pooling methods respectively, resulting in the pooled features Fi under max pooling and global average pooling respectively. p (Dimensions are 1*1*4), for the pooled features F after max pooling and global average pooling respectively. p (Dimension 1*1*4) The features are fused along the channel dimension to obtain fused feature F1 (dimension 1×1×4). Subsequently, fused feature F1 is passed through a fully connected layer to obtain the weights ω of the 4 pixel channels. L (Dimensions 1×1×4).

[0190] Step 423: Process the corresponding pixel channels based on the photosensitive weights of each of the N pixel channels to obtain the photosensitive guide map.

[0191] According to a feasible implementation method, the above input features F0 and ω L Multiplying them together yields the photoguided image G. L For example, when the RAW image is an RGGB image, the input features F0 (dimension (H / 2)*(W / 2)*4) and the weights ω of the 4 channels are... L Multiplying by (dimensions 1×1×4) yields the photoguided image G. L (Dimensions are (H / 2)*(W / 2)*4).

[0192] According to a feasible implementation method, the above input features F0 and ω L After multiplication, summation is performed along the channel dimension to obtain the photosensitive guidance map G. L .

[0193] For example, such as Figure 6b As shown, when the RAW image is an RGGB image, the input features F0 (dimension (H / 2)*(W / 2)*4) and the weights ω of the 4 channels are... L Multiplying (dimensions 1×1×4) and summing them along the channel dimension yields the photosensitive guidance map G. L (Dimensions are (H / 2)*(W / 2)*1).

[0194] Further, based on the steps 421 to 423 shown above Figure 5b In the embodiment of the present application, the step 431 can specifically include the following steps based on the steps 421 to 423 shown above and the step 4311:

[0195] The step 4311 obtains N position encoding maps corresponding to N pixel channels respectively based on the intermediate images of the N pixel channels respectively, and the position encoding map indicates the position number of the corresponding pixel channel and the arrangement of the corresponding pixel channel at the target pixel of the color filter matrix.

[0196] According to a possible implementation, the intermediate images of the N pixel channels are position encoded to obtain a position encoding map M0, which includes N position encoding maps corresponding to the N pixel channels respectively.

[0197] In one example, the N pixel channels are position encoded to obtain position encoding values of the N pixel channels respectively, such as 1, 2, 3, …, N, then for each pixel channel of the N pixel channels, the intermediate image of the pixel channel is arranged according to the pixel pattern of the pixel channel, and the pixel value of the target pixel of the pixel channel is replaced by the corresponding position encoding value, and the pixel values of the target pixels of the pixel channels belonging to other pixel channels are set to 0 to obtain the position encoding map.

[0198] For example, as shown in the following table, when the RAW image is an RGGB image, the intermediate images of different pixel channels in the input feature F0 (dimension (H / 2)*(W / 2)*4) are respectively pattern encoded to obtain a pattern encoding map (dimension (H / 2)*(W / 2)*4), and the position encoding of the pattern encoding map can obtain the position encoding maps M0 (H / 2×W / 2 Figure 6b

[0199] 4) of the 4 pixel channels. Assuming that the position encoding values of the pixel channels R, G, B, and W are 1, 2, 3, and 4 respectively, the pixel values in the position encoding map of the R channel are composed of 1 and 0 (indicating the pixel values of the target pixels of other pixel channels except the R channel), and the G channel, the B channel, and the W channel are similar and will not be described again.

[0200] Further, based on the steps 421 to 423 shown above Figure 5b In the embodiment of the present application, the step 432 can specifically include the following steps based on the steps 421 to 423 shown above and the step 4311:

[0201] The step 4321 processes the light sensitivity weights corresponding to the N pixel channels respectively through the second weight extraction module to determine the position weights corresponding to the N pixel channels respectively.

[0202] Here, the position weights can be more accurately determined by considering the light sensitivity weights.​

[0203] According to one feasible implementation, the second weight extraction module can be a fully connected layer, specifically, the weights ω of the N pixel channels. L (Dimension 1×1×N), after passing through multiple fully connected layers, we obtain the position weights ω of N pixel channels. P (Dimensions 1×1×N).

[0204] For example, such as Figure 6b As shown, when the RAW image is an RGGB image, the weights ω for the four channels are... L (Dimension 1×1×4) After passing through multiple fully connected layers, the position weights ω are obtained. P (Dimensions 1×1×4).

[0205] Furthermore, in step 433, the aforementioned location encoding map M0 and spatial weight ω are... P Multiplying them together yields the spatial guidance graph G. p Here, the spatial guidance graph G P and photoguided map G L They have the same dimensions.

[0206] For example, when the RAW image is an RGGB image, the positional encoding map M0 (dimensions H / 2×W / 2×4) and positional weights ω for the 4-pixel channels are... P Multiplying by (dimensions 1×1×4) yields the spatial guidance graph G. P (Dimensions are (H / 2)*(W / 2)*4).

[0207] According to a feasible implementation, the location coding map M0 and the spatial weight ω are... P After multiplying, summing along the channel dimension yields the spatial guidance graph G. P .

[0208] For example, such as Figure 6b As shown, when the RAW image is an RGGB image, the positional encoding map M0 (dimensions H / 2×W / 2×4) and positional weights ω for the 4-pixel channels are... P Multiplying (dimensions 1×1×4) and then adding them together, we get the spatial guidance graph G. P (Dimensions H / 2 × W / 2 × 1).

[0209] It is worth noting that the photosensitive guide map G in the embodiments of this application... L and spatial guidance diagram G P The first and second weight extraction modules are adaptively generated. If the first and second weight extraction modules are trained with RAW images from different image sensors, then the first and second weight extraction modules can extract the photosensitive guide map G of the RAW images from different image sensors.L and spatial guidance diagram G P .

[0210] According to one feasible implementation, the image restoration network includes M target layers, and the guiding features include guiding feature maps corresponding to each of the m target layers among the M target layers, where m is a positive integer greater than or equal to 2 and less than or equal to M. It is worth noting that the guiding features provided in this embodiment can include guiding feature maps of any and more of the M target layers, and can be used to correct any position in the image restoration network. Here, the target layer indicates the layer that processes the image using a windowing approach; for example, it can be a convolutional layer or a pooling layer.

[0211] like Figure 8 As shown above, in the above Figure 4 Based on the illustrated embodiment, in this embodiment of the application, step 450 may specifically include the following:

[0212] Step 451: During the image restoration process of the image restoration network on the RAW image, when passing through the i-th target layer of m target layers, based on the guiding feature map corresponding to the i-th target layer, the different window data of the target window of the i-th target layer during the sliding window process of the input image are processed; where the input image is the image of the i-th target layer.

[0213] Here, when the target layer is a convolutional layer, the target window is the convolutional kernel; when the target layer is a pooling layer, the target window is the pooling kernel. In some possible implementations, the i-th target layer can also be called the first target layer.

[0214] Figure 9 As shown Figure 4 The flowchart of step 451 in step 450 of the embodiment shown is illustrated.

[0215] like Figure 9 As shown above, in the above Figure 4 Based on the illustrated embodiment, in this embodiment of the application, step 451 may specifically include the following steps:

[0216] Step 4511: During the sliding window process of the target window of the i-th target layer on the input image, determine the current window data of the i-th target layer; wherein, the window data is the data of the current position of the target window on the input image, and the input image and the guiding feature map corresponding to the i-th target layer are adapted in terms of planar size.

[0217] In some possible cases, the input image and the guiding feature map corresponding to the i-th target layer may be the same or different in planar dimensions, but they are usually the same. This application's embodiments use the example of the input image and the guiding feature map corresponding to the i-th target layer being adapted in planar dimensions for illustration.

[0218] Here, for ease of description, the input image is denoted as Fi, with a size of h*w*c1; where h represents the height of the input image, w represents the width of the input image, and c represents the number of sub-images in the input image, i.e., the number of channels. The i-th convolutional layer is denoted as Ci, and the target window of Ci is denoted as Ki, with a size of (k1*k2*c1); where k1 represents the length of the target window, k2 represents the width of the target window, and c1 represents the number of target windows, i.e., the number of channels of the target window, which is consistent with the number of channels of the input image. The window data is denoted as fn, which can be understood as the spatial data obtained by the target window Ki during movement on the input image Fi. The size of the window data fn is the same as that of the target window Ki, both being k1*k2*c1. Here, n in fn represents the number of pixel points, which is equal to k1*k2. The guided feature map corresponding to Ci is denoted as Gi, with a size of h*w*c2. Wherein, the guided feature map Gi and the input image Fi are the same in the plane size, both being h*w, but the channel dimension can be different or the same, in other words, c1 and c2 can be the same or different, and the embodiments of the present application do not make specific limitation thereon.

[0219] Step 4512, determining first guided data corresponding to the window data; wherein the first guided data is data in the guided feature map corresponding to the first target layer.

[0220] For ease of description, it is assumed that the first guided data is denoted as Gn, with a size of k1*k2*c2. n in Gn represents the number of pixel points, which is the same as and corresponds to the number of pixel points of the window data. Since the first guided data Gn and the target window Ci are the same in the plane size, after aligning the input image and the guided feature map Gi, the data in the space of the guided feature map Gi corresponding to the position of the target window Ci on the input image is taken as the first guided data Gn.

[0221] Step 4513, processing the window data based on the first guided data.

[0222] Specifically, based on the first guided data, a correction value corresponding to the window data is determined; wherein the correction value indicates the correlation between different pixel points in the window data; and the window data is processed based on the correction value.

[0223] Here, the correlation can be understood as the correlation of the RAW domain. The RAW domain refers to the process from DPC (bad pixel correction), BLC (black level correction) to demosaic (noise reduction). The bad pixel is caused by the chip manufacturing process and other problems. The bad pixel refers to the brightness or color that is very different from other pixels around it. The black level refers to the signal level corresponding to the image data of 0. Noise often appears as an isolated pixel point or pixel block on the image, which causes strong visual effects.

[0224] According to a possible implementation, the correction value corresponding to the window data is determined based on the first guide data in the following manner.

[0225] A target pixel point in the first guide data is determined. The target pixel point indicates the position of the pixel point obtained after processing the window data in the input image. For each pixel point in the first guide data, the similarity between the pixel point and the target pixel point is calculated based on the data of the pixel point and the target pixel point in the first guide data, and the similarity is taken as the correction value of the position point of the pixel point in the window data.

[0226] For example, assuming that the size of the target window Ki is 3*3*c1, the target pixel point is the center point in the 3*3 matrix. Assuming that the size of the target window Ki is 2*2*c1, the target pixel point is the right lower pixel point in the 2*2 matrix.

[0227] It is worth noting that the target window Ki shares the correction value in the channel dimension. Specifically, the data of n pixel points in the window data fn is represented as f1, f2, …, fn; the correction value includes the correction value C1, C2, …, Cn corresponding to f1, f2, …, fn respectively; correspondingly, f1*c1 in the window data fn is corrected to C11*f1*c1, and f2, …, fn are similar and will not be described again. Then, the corrected window data is processed. For example, when the target window Ki is a convolution kernel, the corrected window data is subjected to a convolution operation; when the target layer is a pooling layer, if the maximum pooling method is used, the maximum value in the corrected window data can be selected; if the average pooling method is used, the corrected window data can be averaged.

[0228] Taking the convolution operation as an example, first, the neighborhood similarity a i,j (j = 1, 2, …, n) is calculated.

[0229] a i,j = σ((g i ) T ·g j )

[0230] Where σ(·) is the sigmoid activation function, i represents the position index of k1*k2 (usually the center position), and j is the position index of the corresponding neighborhood feature, which is 1, 2, ..., n.

[0231] via a i,j For the current window data f n Perform a convolution operation to obtain the corrected feature f′ n :

[0232]

[0233] Subsequently, f′ n This represents the result of performing a convolution operation on the window data.

[0234] For example, suppose the target layer is a convolutional layer, such as Figure 7c As shown, the size of the convolution kernel Ki is 3*3*7, and the size of the input image Fi is 5*5*5. During the sliding window process of the convolution kernel Ki on the input image Fi, the window data fn is determined to be 3*3*7, and the first guiding data Gn is 3*3*5. The pixel G22 at the center of the first guiding data Gn is taken as the target pixel. The similarity between the target pixel and each pixel in the first guiding data Gn is calculated, resulting in a weighted correction image of 3*3*1. C11 in the weighted correction image represents the fusion result between the space of the target pixel G22 and the space of pixel G11. C12, ..., ... C33 is similar and will not be elaborated further. Next, the spaces containing position K11 in the convolution kernel Ki and position P11 in the convolution space fn are multiplied along the channel dimension to obtain a calculation result of size 1*1*7. Based on C11, the value of each channel in the calculation result is corrected to obtain the final calculation result p11 (size 1*1*7). The calculation results p12, ..., p33 are calculated in the same way. Then, the calculation results of p11, ..., p33 are fused along the channel dimension, i.e., averaged, to obtain the processing result of window data fn (scale 1*1*7). The processing results of each window data during the sliding windowing process of convolution kernel Ki on the input image Fi are processed in the same way. Since the first guiding data is different, the processing of different window data is different.

[0235] It is worth noting that in the embodiments of the present application, the guide data of different window data is different, so the processing method of different window data is different. In addition, in the foregoing description, the input image and the guide feature map corresponding to the i-th target layer are adapted in the plane size, but in actual application, the input image and the guide feature map corresponding to the i-th target layer can not be adapted in the plane size, so that when the first guide data is determined, the number of data in the first guide data can be higher or lower than the number of data in the window data. In this way, when the first guide data processes the window data, the first guide data can be up-sampled or down-sampled to obtain the first guide data that is adapted to the window data.

[0236] It should be noted that in the embodiments of the present application, the target layer is used to filter out more effective information for image recovery, which can be understood as "refining" of effective information. When filtering effective information, the correlation between the pixels in the region is calculated considering the difference in light sensing characteristics and spatial characteristics of different positions in the region. Since the guide is based on the light sensing characteristics and spatial characteristics, the information is determined to be effective or not by the light sensing characteristics and spatial characteristics, so more information that meets the light sensing characteristics and spatial characteristics can be retained. In addition, the correlation of different regions is different, so the embodiments of the present application extract effective information in different ways through different window data, so that different positions are processed in different ways, so that more details and textures of the image can be retained and the generation of artifacts can be reduced.

[0237] In the present scheme, the processing results of the window at different positions are corrected by the guide feature map, so that different pixel points are processed in different ways. At the same time, the correlation between different pixel points in the window data is considered in the correction process, so that the image recovery network can retain more image details and textures, reduce the generation of artifacts, and ensure the image recovery quality.

[0238] In one example, the M target layers in the image recovery network can form an image encoding network and an image decoding network; wherein the m target layers for guiding are distributed in the image encoding network and / or the image decoding network. The decoding network in the image recovery network provided by the embodiments of the present application generates deep, semantic, and coarse-grained feature maps, which are combined with the shallow, low-level, and fine-grained feature maps generated by the image encoding network to obtain better segmentation capability and improve the image recovery effect.

[0239] In actual application, the image encoding network includes n encoding layers and m decoding layers.

[0240] Here, the encoding layer can extract shallow, low-level, and fine-grained features. In the embodiments of the present application, the encoding layer is composed of at least one convolutional layer, and when there are multiple convolutional layers, the multiple convolutional layers can be connected in series.

[0241] Here, the encoding layers can extract deep, semantic, coarse-grained features. In the embodiments of the present application, the decoding layers are composed of at least one convolutional layer, and when there are multiple convolutional layers, the multiple convolutional layers can be connected in series.

[0242] Optionally, the n encoding layers are connected in series, and the m decoding layers are connected in series.

[0243] In one example, for the i-th encoding layer in the n encoding layers connected in series, the input image (referred to as an input image for ease of description and distinction) is the down-sampled result of the output of the connected previous encoding layer (referred to as the i-1-th encoding layer for ease of description and distinction).

[0244] Further, n is greater than m, and each decoding layer is connected to at least one encoding layer. For example, n = m + 1, the last encoding layer in the n encoding layers connected in series is connected to the first decoding layer in the m decoding layers connected in series, and the output of the i-th encoding layer is used as the input of the m-i+1-th decoding layer.

[0245] In addition, for the j-th decoding layer in the m decoding layers connected in series, the input image is the result of splicing the processed feature map of the output of all the connected encoding layers and the up-sampled output of the previous decoding layer (referred to as the j-1-th decoding layer for ease of description and distinction). It is worth noting that the output of the decoding layer and the up-sampled output need to be spliced as the input of the decoding layer, the purpose is to fuse the feature information, so that the information of the deep layer and the shallow layer is fused, and when splicing, not only the picture size needs to be consistent, but also the dimension (channels) of the feature needs to be consistent, so that the splicing can be performed.

[0246] In one example, the image restoration network can be a Unet network. The Unet network can be mainly divided into two parts, a feature extraction network and an up-sampling network. Here, the image encoding network can be the feature extraction network, and the up-sampling network can be understood as the image decoding network.

[0247] The feature extraction network is a shrinking network, which includes 5 shallow feature extraction modules (encoding layers) and 4 down-sampling modules (for implementing down-sampling) to reduce the picture size, and in the process of continuous down-sampling, multi-scale feature maps can be extracted, i.e., the scale of the feature map has multiple scales. The shallow feature extraction module is composed of two convolutional layers connected in series. The down-sampling module can be a pooling layer. The shallow feature extraction module and the down-sampling module are arranged alternately.

[0248] For example, suppose the input image size is 572x572. After passing through two convolutional layers (with a kernel size of 3x3) in the shallow feature extraction module, the size of the feature map changes from 572*572 to 570*570 to 568*568. Then, after passing through the downsampling module (let's say Maxpool with a kernel size of 2x2), the image size becomes 284. This is a complete downsampling. Therefore, during the downsampling process, the number of channels doubles, for example, from 64 to 128. The next three downsampling processes are similar.

[0249] The upsampling network, also called an expanded network, enlarges the image size and extracts deeper information. It uses five deep feature extraction modules (decoding layers) and four upsampling modules. The deep feature extraction modules consist of two concatenated convolutional layers. The upsampling modules can be deconvolutional layers or interpolation layers, used to upsample the outputs of either the deep or shallow feature extraction modules. During upsampling, the number of image channels is halved, the opposite of the change in the number of channels in the shallow feature extraction. The upsampling process incorporates information from the shallower layers on the left, essentially stitching together the features from the left side.

[0250] like Figure 7d As shown in the embodiment of this application, an image restoration network under a U-net network is illustrated. n=5, m=4, in other words, the image restoration network includes 5 decoding layers and 4 encoding layers.

[0251] The first encoding layer can include a first convolutional layer Conv(k3n64s1) and a second convolutional layer Conv(k3n64s1) in series. Here, k represents the kernel size, n represents the number of channels in the convolutional feature map, and s represents the stride. That is, the kernel size of convolutional layer Conv(k3n64s1) is 3, the number of channels in the convolutional feature map is 64, and the stride is 1.

[0252] The second coding layer may include a first convolutional layer Conv(k3n128s1) and a second convolutional layer Conv(k3n128s1).

[0253] The input to the first convolutional layer Conv(k3n128s1) is the downsampled output of the second convolutional layer Conv(k3n64s1) in the first coding layer. Here, the size of the output of the second convolutional layer Conv(k3n64s1) in the first coding layer is reduced by half after downsampling.

[0254] The third coding layer may include a first convolutional layer Conv(k3n256s1) and a second convolutional layer Conv(k3n256s1).

[0255] The input of the first convolutional layer Conv(k3n256s1) is the result of the output of the second convolutional layer Conv(k3n128s1) in the second encoding layer after down-sampling, where the size of the output of the second convolutional layer Conv(k3n128s1) in the second encoding layer is reduced by half after down-sampling.

[0256] For the fourth encoding layer, the first convolutional layer Conv(k3n512s1) and the second convolutional layer Conv(k3n512s1) can be included.

[0257] The input of the first convolutional layer Conv(k3n512s1) is the result of the output of the second convolutional layer Conv(k3n256s1) in the third encoding layer after down-sampling, where the size of the output of the second convolutional layer Conv(k3n256s1) in the third encoding layer is reduced by half after down-sampling.

[0258] For the fifth encoding layer, the first convolutional layer Conv(k3n1024s1) and the second convolutional layer Conv(k3n1024s1) can be included.

[0259] The input of the first convolutional layer Conv(k3n512s1) is the result of the output of the second convolutional layer Conv(k3n512s1) in the fourth encoding layer after down-sampling, where the size of the output of the second convolutional layer Conv(k3n512s1) in the fourth encoding layer is reduced by half after down-sampling.

[0260] For the first decoding layer, the first convolutional layer Conv(k3n512s1) and the second convolutional layer Conv(k3n512s1) can be included in series.

[0261] The input of the first convolutional layer Conv(k3n512s1) is the result of the output of the second convolutional layer Conv(k3n512s1) in the fifth encoding layer after up-sampling, and the result of the output of the second convolutional layer Conv(k3n512s1) in the fourth encoding layer after size adjustment; where the size of the output of the second convolutional layer Conv(k3n512s1) in the fifth encoding layer is doubled after up-sampling, and the output of the second convolutional layer Conv(k3n512s1) in the fourth encoding layer after size adjustment and the result of the output of the second convolutional layer Conv(k3n512s1) in the fifth encoding layer after up-sampling are adapted in channel (all 512) and size.

[0262] For the second decoding layer, the first convolutional layer Conv(k3n256s1) and the second convolutional layer(k3n256s1) can be included in series.

[0263] The input of the first convolutional layer Conv(k3n256s1) is the result of splicing the result of up-sampling the output of the second convolutional layer Conv(k3n512s1) in the first decoding layer with the result of adjusting the size of the output of the second convolutional layer Conv(k3n256s1) in the third encoding layer; wherein the output of the second convolutional layer Conv(k3n512s1) in the first decoding layer is doubled in size after up-sampling, the output of the second convolutional layer Conv(k3n256s1) in the third encoding layer is adjusted in size, and the result of up-sampling the output of the second convolutional layer Conv(k3n512s1) in the first decoding layer is adapted in the channel (all 256) and size.

[0264] For the third decoding layer, the first convolutional layer Conv(k3n128s1) and the second convolutional layer Conv(k3n128s1) in series can be included.

[0265] The input of the first convolutional layer Conv(k3n128s1) is the result of splicing the result of up-sampling the output of the second convolutional layer Conv(k3n256s1) in the second decoding layer with the result of adjusting the size of the output of the second convolutional layer Conv(k3n128s1) in the second encoding layer; wherein the output of the second convolutional layer Conv(k3n256s1) in the second decoding layer is doubled in size after up-sampling, the output of the second convolutional layer Conv(k3n128s1) in the second encoding layer is adjusted in size, and the result of up-sampling the output of the second convolutional layer Conv(k3n256s1) in the second decoding layer is adapted in the channel (all 128) and size.

[0266] For the fourth decoding layer, the first convolutional layer Conv(k3n64s1) and the second convolutional layer in series can be included.

[0267] The input of the first convolutional layer Conv(k3n64s1) is the result of splicing the result of up-sampling the output of the second convolutional layer Conv(k3n128s1) in the third decoding layer with the result of adjusting the size of the output of the second convolutional layer Conv(k3n64s1) in the first encoding layer; wherein the output of the second convolutional layer Conv(k3n128s1) in the third decoding layer is doubled in size after up-sampling, the output of the second convolutional layer Conv(k3n64s1) in the first encoding layer is adjusted in size, and the result of up-sampling the output of the second convolutional layer Conv(k3n128s1) in the third decoding layer is adapted in the channel (all 64) and size.

[0268] It is worth noting that in the embodiments of the present application, the first convolutional layer and the second convolutional layer will be processed by the activation function after the subsequent processing. In addition, the down-sampling can be a pooling layer, and the up-sampling can be a de-convolutional layer. The up-sampling and the down-sampling can be guided by the guided feature map.

[0269] It should be understood that the embodiments of the present application only provide an exemplary structure of the image restoration network, and other structures can also be used, for example, the number of convolutional layers and activation function layers can not be 2, and the numbers of k, n, and s in the convolutional layer are optional. Specifically, corresponding to different sizes of RAW images, the structure of the image restoration network also needs to be adjusted accordingly. For example, the number of encoding layers and / or decoding layers will be different.

[0270] Based on the embodiments shown in the above Figure 4 and Figure 8 In the embodiments of the present application, step 440 specifically includes the following content:

[0271] The guided feature extraction module is used to extract guided features from the light sensing guide map and the spatial guide map, and determine n guided feature maps output by the guided feature extraction module.

[0272] According to a possible implementation manner, the guided feature extraction module and the structure of the image restoration network are adapted, and the network can also be of an encoding-decoding type.

[0273] Specifically, the guided feature extraction module can include Q target layers, and the Q target layers form x guided feature encoding units and y guided feature decoding units. Here, the output of each of the m target layers in the Q target layers can be used as a guided feature map.

[0274] It is worth noting that as long as the guided feature map and the input map of the target layer are the same in the plane size, the embodiments of the present application can modify the weights of any convolutional layer in the image restoration network.

[0275] Optionally, the number of guided feature encoding units is the same as the number of image encoding units in the image restoration network. Meanwhile, the number of target layers in the guided feature encoding unit is equal to the number of target layers in the image encoding unit, and the guided feature map output by the target layer of the guided feature encoding unit and the input map of the target layer of the image encoding unit are the same in the plane size.

[0276] As Figure 6b shown, a specific example 1 of an image restoration architecture is given below. The image restoration architecture includes a guide map extraction module, a guided feature extraction network, and an image restoration network. The following takes an example of performing image restoration on a RAW image in RGGB format.

[0277] First, receive the input RAW image I. input (Dimensions H×W×1), after pixel shift operation, the intermediate image F0 (dimensions H / 2×W / 2×4) is obtained.

[0278] Guided graph extraction module: On one hand, the obtained F0 is processed by Maxpool and Avgpool operations respectively to obtain two pooled features F. M (dimensions 1×1×4) and F A (Dimension 1×1×4), concatenate the channels to obtain the fused feature F1 (Dimension 1×1×4), flatten it, and then pass it through multiple fully connected layers to obtain the guide graph weights ω. L (Dimension 1×1×4); Input features F0 and ω L Multiplying and summing the results yields the photoguided image G. L (Dimension H / 2 × W / 2 × 1); On the other hand, the above I input Position encoding is performed to obtain a 4-pixel channel position encoded map M0 (dimensions H / 2×W / 2×4). The weights ω of the above guide map are then applied. L After passing through multiple fully connected layers, the spatial weights ω are obtained. P (Dimension 1×1×4); Input features M0 and ω P Multiplying and summing the results yields the spatial guidance graph G. P (Dimension h / 2 × W / 2 × 1). See Figure 6 and steps 421 to 423 and 431 to 433 above for details. The guide graph extraction module consists of a first weight extraction module and a second weight extraction module.

[0279] Guided feature extraction module: The input is the photosensitive guide image G mentioned above. L and spatial guidance diagram G P The guided input features G0 (dimensions H / 2 × W / 2 × 2) are obtained by concatenating the channels. After passing through M convolutional layers, the guided features g of the M convolutional layers are extracted. n , n∈{1,2,...,M}.

[0280] Image Restoration Network 300: The input consists of the above-mentioned input features F0 (dimensions H / 2×W / 2×4) and g n After passing through M convolutional layers, the restored image I is obtained. out (Dimensions H×W×1 or H×W×3).

[0281] The image restoration network employs adaptive convolution. For the adaptive convolution process of the i-th convolutional layer out of M convolutional layers, please refer to [link / reference needed]. Figure 7d and against Figure 7d The description will not be repeated here.

[0282] The training process of the guidance map extraction module, the guidance feature extraction module, and the image restoration network 300 is described below. For ease of description, the guidance map extraction module, the guidance feature extraction module, and the image restoration network 300 are referred to as a restoration model.

[0283] First, a picture set is constructed, which includes a plurality of RAW images and labels of the plurality of RAW images. The plurality of RAW images can come from different image sensors, and the different image sensors have different color filter matrices in front.

[0284] In a possible case, the label can be an image after the RAW image is restored.

[0285] The restoration model can then be trained by the difference between the RAW image and the restored image output by the restoration model when the RAW image is input into the restoration model.

[0286] In a possible case, the label can be the position and / or category of the target in the RAW image. Then, when training the restoration model, the restored image output by the restoration model needs to be input into a target detection model to obtain a detection result, which is a label adaptation, for example, can include the positions of a plurality of target boxes and the categories to which each target box belongs. Subsequently, when training the restoration model, the restoration model and the target detection model are trained based on the difference between the RAW image and the detection result of the RAW image input into the restoration model and then input into the target detection model.

[0287] It is worth noting that the light-sensing guidance map G L and the spatial guidance map G P are adaptively generated by the guidance map extraction module. If the parameters of the guidance map extraction module are trained using RAW images of different image sensors, the guidance map extraction module can extract the light-sensing guidance map G L and the spatial guidance map G P of the RAW images of different image sensors.

[0288] In a specific example 1, by considering the guidance features of the spatial characteristics and the light-sensing characteristics of the RAW image, the image restoration network is assisted to restore the RAW image, so that the image restoration network can perform different processing on different feature points in the features extracted at different positions, implement different processing on target pixels of different pixel arrangements of the N pixel channels, so that the image restoration network extracts information meeting the light-sensing characteristics and the spatial characteristics, improves the details and textures after the RAW image is restored, reduces artifacts, and ensures the image quality.

[0289] For example, Figure 6aAs shown, another specific example of an image restoration architecture, Example 2, is given below. This image restoration architecture includes: a guided image extraction module, a guided feature extraction network, and an image restoration network. A specific example of image restoration using RGGB format RAW images is provided below.

[0290] First, receive the input RAW image I. input (dimensions H×W×1) and the photoguide map G obtained by the brightness sensor L (Dimensions H×W×1).

[0291] For RAW image I input Position encoding is performed on the N pixel channels in the RAW image, and the RAW image I... input The pixel values ​​of N pixel channels in the image are replaced with the corresponding position codes to obtain the spatial guidance map G. P (Dimensions H×W×1).

[0292] Guided feature extraction module: The input is the photosensitive guide image G mentioned above. L and spatial guidance diagram G P The guided input features G0 (dimensions H=W=2) are obtained by concatenating the channels. After passing through M convolutional layers, the guided features g of the M convolutional layers are extracted. n , n∈{1,2,...,M}.

[0293] Image Restoration Network 300: Input is the RAW image I mentioned above. input and g n After passing through M convolutional layers, the restored image I is obtained. out (Dimensions H×W×1 or H×W×3).

[0294] The image restoration network employs adaptive convolution. For the adaptive convolution process of the i-th convolutional layer out of M convolutional layers, please refer to [link / reference needed]. Figure 7d and against Figure 7d The description will not be repeated here.

[0295] The training process of the guided feature extraction module and the image restoration network 300 is described below. The guided feature extraction module and the image restoration network 300 can be used as a restoration model; see the specific example 1 above for the training process of the restoration model.

[0296] For the beneficial effects of specific example 2, please refer to specific example 1. In addition, compared with specific example 1, since the photoguided image G is directly acquired by the brightness sensor, L This ensures the photoguided image G L Its reference value.

[0297] like Figure 6aAs shown, another specific example of an image restoration architecture, number 3, is given below. This image restoration architecture includes: a guided image extraction module, a guided feature extraction network, and an image restoration network. A specific example of image restoration using RGGB format RAW images is provided below.

[0298] First, receive the input RAW image I. input (dimensions H×W×1) and the photoguide map G obtained by the brightness sensor L (Dimensions H×W×1).

[0299] For RAW image I input Position encoding is performed on the N pixel channels in the RAW image, and the RAW image I... input The pixel values ​​of the four pixel channels in the image are replaced with the corresponding position codes to obtain the encoded image RAW. code (Dimensions H×W×1);

[0300] The guide map extraction module takes the aforementioned photosensitive guide map G as input. L and encoding graph RAW code The above photosensitive guidance map G L After passing through multiple fully connected layers, the spatial weights ω are obtained. P (Dimensions 1×1×4); Encode the RAW image code The four positional codes and their corresponding positions in ω P Multiplying the values ​​in the graphs yields the spatial guidance graph G. P (Dimensions H×W×1).

[0301] Guided feature extraction module: The input is the photosensitive guide image G mentioned above. L and spatial guidance diagram G P The guided input feature G0 (dimensions H×W×2) is obtained by concatenating the channels. After passing through M convolutional layers, the guided feature g of the M convolutional layers is extracted. n , n∈{1,2,…,M}.

[0302] Image Restoration Network 300: Input is the RAW image I mentioned above. input and g n After passing through M convolutional layers, the restored image I is obtained. out (Dimension H = W = 1 or H × W × 3).

[0303] The image restoration network employs adaptive convolution. For the adaptive convolution process of the i-th convolutional layer out of M convolutional layers, please refer to [link / reference needed]. Figure 7d and against Figure 7d The description will not be repeated here.

[0304] The training process of the guidance map extraction module, the guidance feature extraction module and the image restoration network 300 is described below. The guidance map extraction module, the guidance feature extraction module and the image restoration network 300 can be used as a restoration model. For details, refer to the training process of the restoration model in the foregoing specific example 1.

[0305] It is worth noting that the spatial guidance map G P The spatial guidance map G P .

[0306] For details, refer to specific example 2. In addition, compared with specific example 2, the light sensor is used to directly collect the light sensing guidance map G L The reference value of the spatial guidance map G P is ensured, and thus the image restoration effect is ensured.

[0307] Based on the same concept as the method embodiments of the present application, the image restoration device is also provided in the embodiments of the present application. The image restoration device includes a plurality of modules, each module is used to execute each step in the image restoration method provided by the embodiments of the present application, and the division of the modules is not limited herein. Those skilled in the art can clearly understand that in actual application, each step in the image restoration method provided by the embodiments of the present application can be completed by different modules according to needs, that is, the internal structure of the device is divided into different modules to complete all or part of the functions described above. Each module in the embodiments can be integrated in one processing unit, or each unit can exist physically, or two or more modules can be integrated in one unit. The integrated unit can be realized in the form of hardware or in the form of a software function unit. In addition, the specific names of each module are only used for easy distinction, and do not limit the protection scope of the present application. For the specific working process of the modules in the device, refer to the corresponding process in the foregoing method embodiments, which will not be described herein.

[0308] For example, the image restoration device is used to execute the personal information identification method provided by the embodiments of the present application, Figure 10 is a structural schematic diagram of the image restoration device provided by the embodiments of the present application. As Figure 10 shown, the image restoration device provided by the embodiments of the present application includes:

[0309] The image acquisition module 1001 is configured to acquire a RAW image captured by an image sensor on a target scene; wherein the image sensor acquires the RAW image through a front color filter matrix, the color filter matrix is arranged in a target pixel unit, the target pixel unit is arranged by N target pixels, target pixels at the same position of each target pixel unit in the color filter matrix form a pixel channel, and N pixel channels are formed in total.

[0310] The first guide map determination module 1002 is configured to determine a light sensing guide map corresponding to the RAW image, wherein the light sensing guide map indicates brightness information of the RAW image.

[0311] The second guide map determination module 1003 is configured to determine a spatial guide map corresponding to the RAW image; wherein the spatial guide map indicates a positional relationship of the N pixel channels.

[0312] The feature extraction module 1004 is configured to perform feature extraction to determine a guide feature of an image restoration network based on the light sensing guide map and the spatial guide map.

[0313] The image restoration module 1005 is configured to perform image restoration on the RAW image based on the guide feature and according to the image restoration network, to obtain a restored image corresponding to the RAW image.

[0314] In a possible implementation, the guide feature indicates extracted light sensing information and information of the RAW image obtained by the positional relationship of the N pixel channels, which can reflect the light sensing difference of the N pixel channels and the difference in pixel arrangement of the target pixels of the N pixel channels; wherein the pixel arrangement indicates the distribution of the target pixels belonging to the same pixel channel around the target pixels.

[0315] In a possible implementation, the image restoration network considers the light sensing difference of the N target pixels of the RAW image, and adopts different processing for the target pixels of the N pixel channels with different pixel arrangements to realize image restoration; wherein the pixel arrangement indicates the distribution of the target pixels belonging to the same pixel channel around the target pixels.

[0316] In a possible implementation, the image restoration module is configured to process at least part of the features extracted by the image restoration network based on the guide feature.

[0317] In a possible implementation, the number of target layers in the image restoration network is M, the guide features include m target layers of the M target layers each corresponding to a guide feature map, M is a positive integer greater than or equal to 2, and m is a positive integer less than or equal to M; the target layer is a convolutional layer or a pooling layer; the image restoration module is configured to, in the process of image restoration of the RAW image by the image restoration network, when a first target layer of the m target layers is passed, process window data of a target window of the first target layer on an input graph based on a guide feature map corresponding to the first target layer; the input graph is a graph input to the first target layer, and the first target layer is any target layer of the m target layers.

[0318] In one example, the image restoration module includes a first space determination unit, a second space determination unit, and a processing unit.

[0319] The first space determination unit is configured to determine window data of a target window of the first target layer; the window data is data of a position of the target window on the input graph.

[0320] The second space determination unit is configured to determine first guide data corresponding to the window data; the first guide data is data in the guide feature map corresponding to the first target layer.

[0321] The processing unit is configured to process the window data based on the first guide data.

[0322] In a possible case, the first guide data corresponding to the window data of the target window at different positions on the input graph is different.

[0323] In a possible case, the input graph and the guide feature map corresponding to the first target layer are the same in terms of plane size, and the first guide data is data in the guide feature map corresponding to the first target layer that fits the position of the target window in the input graph.

[0324] In a possible case, the processing unit is configured to determine a correction value corresponding to the window data based on the first guide data; the correction value indicates a correlation between different pixel points in the window data; and the window data is processed based on the correction value.

[0325] Optionally, the processing unit is configured to determine a target pixel point in the first guide data, wherein the target pixel point indicates a position of a pixel point obtained after processing the window data in the input image; and for each pixel point in the first guide data, calculate a similarity between the pixel point and the target pixel point based on data of the pixel point and the target pixel point in the first guide data, and take the similarity as a correction value of the pixel point.

[0326] In one example, the image restoration network comprises an image encoding network and an image decoding network; and m target layers are distributed in the image encoding network and / or the image decoding network.

[0327] In one example, the feature extraction module is configured to perform feature extraction on the light-sensing guide image and the spatial guide image by a guide feature extraction module, and determine n guide feature maps output by the guide feature extraction module.

[0328] In a possible implementation, the first guide map determination module comprises a pixel recombination unit, a first weight determination unit and a first guide map determination unit; wherein

[0329] The pixel recombination unit is configured to determine intermediate images corresponding to the N pixel channels in the RAW image respectively;

[0330] The first weight determination unit is configured to determine light-sensing weights corresponding to the N pixel channels respectively based on the intermediate images corresponding to the N pixel channels respectively;

[0331] The first guide map determination unit is configured to process the corresponding pixel channels based on the light-sensing weights corresponding to the N pixel channels respectively to obtain a light-sensing guide map.

[0332] In a possible implementation, the second guide map determination module comprises a position encoding unit, a second weight determination unit and a second guide map determination unit; wherein

[0333] The position encoding unit is configured to obtain position encoding maps corresponding to the N pixel channels respectively based on the RAW image, wherein the position encoding maps indicate position numbers of the corresponding pixel channels;

[0334] The second weight determination unit is configured to determine position weights corresponding to the N pixel channels respectively;

[0335] The second guide map determination unit is configured to correct the position encoding maps corresponding to the N pixel channels respectively based on the position weights corresponding to the N pixel channels respectively to obtain a spatial guide map.

[0336] Here, the above Figure 6aThe spatial guidance map G is determined in the manner of the illustrated implementation modes B1 and B2 P .

[0337] In one example, the position coding map also indicates the arrangement of the corresponding pixel channels in the target pixel of the color filter matrix.

[0338] Here, the spatial guidance map can be determined in the manner of the above steps 4031, 4032 and 403.

[0339] In one possible implementation, the light sensing guidance map is an image obtained by a luminance sensor capturing the target scene.

[0340] In addition to the above method, device and electronic equipment, the embodiments of the present application can also provide a computer program product, which includes computer program instructions, and when the computer program instructions are executed by a processor, the processor executes the steps in the image recovery method of various embodiments of the present application described in the above “method” part of the specification. Wherein the computer program product can be written in one or more program design languages in any combination for computer program code for executing the operations of the embodiments of the present application, the program design languages include object-oriented program design languages such as Java, C++, etc., and conventional procedural program design languages such as “C” language or similar program design languages. Wherein the computer program code can be in the form of source code, object code, executable file or some intermediate form, etc. The computer program code can be completely executed on a user computing device, partially executed on a user device, executed as a separate software package, partially executed on a user computing device and partially on a remote computing device, or completely executed on a remote computing device or server.

[0341] Furthermore, embodiments of the present application can also provide a computer readable storage medium having computer program instructions stored therein, which when executed by a processor, cause the processor to perform the steps of the image recovery method according to various embodiments of the present disclosure described in the above "METHOD" section of the specification. The computer readable storage medium can employ any combination of one or more of readable media. The readable medium can be a readable signal medium or a readable storage medium. The readable storage medium may, for example, include but is not limited to an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the above. More specific examples (a non-exhaustive list) of the readable storage medium include an electrical connection having one or more wires, a portable disc, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. It should be noted that the computer readable medium contains content according to the requirements of the legislation and patent practice in the jurisdiction, and the computer readable medium can be appropriately added or reduced, for example, in some jurisdictions, according to the legislation and patent practice, the computer readable medium does not include electrical carrier signals and telecommunication signals.

[0342] The method steps in the embodiments of the present application can be realized by hardware or by the processor executing software instructions. The software instructions can be composed of corresponding software modules, and the software modules can be stored in a random access memory (RAM), a flash memory, a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically EPROM (EEPROM), a register, a hard disk, a mobile hard disk, a CD-ROM, or any other form of storage medium well known in the art. An exemplary storage medium is coupled to the processor, so that the processor can read information from and write information to the storage medium. Of course, the storage medium can also be an integral part of the processor. The processor and the storage medium can be located in an ASIC.

[0343] In the embodiments described above, all or some of the steps can be implemented by hardware, software, firmware or any combination thereof. When implemented in software, all or some of the steps can be implemented in the form of a computer program product. The computer program product includes one or more computer instructions. When loaded and executed by a computer, all or some of the computer instructions can generate the processes or functions described in the embodiments of the present application. The computer can be a general purpose computer, a special purpose computer, a computer network, or other programmable apparatus. The computer instructions can be stored in a computer readable storage medium or transmitted by a computer readable storage medium. The computer instructions can be transmitted from one website, computer, server or data center to another website, computer, server or data center through a wired (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (such as infrared, wireless, microwave, etc.) manner. The computer readable storage medium can be any available medium accessible by a computer or a data storage device such as a server, data center, etc. integrated with one or more available media.

[0344] In the embodiments described above, the description of each embodiment has its own focus, and the parts not described or recorded in detail in a certain embodiment can be referred to the related description of other embodiments.

[0345] It should be understood that the size of the serial number of each step in the above embodiments does not mean the order of execution, and the execution order of each process should be determined according to its function and inherent logic, and should not constitute any limitation on the implementation process of the embodiments of the present application.

[0346] The basic principles of the present application are described above in combination with specific embodiments, but it should be pointed out that the advantages, advantages, effects mentioned in the present application are only examples and not limitations, and these advantages, advantages, effects cannot be considered as the must-have of each embodiment of the present application. In addition, the specific details of the above disclosure are only for the purpose of example and understanding, and are not limited to the specific details described above.

[0347] The block diagrams of the devices, apparatuses, equipment, systems involved in the present disclosure are only illustrative examples and are not intended to require or imply the connection, arrangement, configuration shown in the block diagram. As those skilled in the art will recognize, these devices, apparatuses, equipment, systems can be connected, arranged, configured in any manner. Words such as "include", "contain", "have" and the like are open-ended words, which mean "including but not limited to", and can be used interchangeably. The words "or" and "and" used herein mean the word "and / or", and can be used interchangeably unless the context clearly indicates otherwise. The word "such as" used herein means the phrase "such as but not limited to", and can be used interchangeably.

[0348] It should also be noted that in the apparatus, device and method of the present disclosure, each component or each step can be decomposed and / or recombined. These decompositions and / or recombinations should be considered as equivalent solutions of the present disclosure.

[0349] The above description has been presented for the purposes of illustration and description. Furthermore, this description is not intended to limit the embodiments of the present disclosure to forms disclosed herein. Although several example aspects and embodiments have been discussed above, those skilled in the art will recognize certain variations, modifications, changes, additions and sub-combinations thereof.

[0350] It can be understood that various numerical numbers involved in the embodiments of the present application are only distinguished for the convenience of description, and are not used to limit the scope of the embodiments of the present application.

Claims

1. An image restoration method, characterized in that, include: A RAW image of a target scene is acquired by an image sensor; wherein the image sensor acquires the RAW image through a front-end color filter matrix, the color filter matrix is ​​arranged in target pixel units, the target pixel unit is arranged in N target pixels, and the target pixels at the same position in each target pixel unit in the color filter matrix form a pixel channel, and a total of N pixel channels are formed. Determine the photosensitive guide map corresponding to the RAW image, wherein the photosensitive guide map indicates the brightness information of the RAW image; Determine the spatial guide map corresponding to the RAW image; wherein the spatial guide map indicates the positional relationship of the N pixel channels; Based on the photosensitive guidance map and the spatial guidance map, feature extraction is performed to determine the guidance features of the image restoration network; Based on the guiding features, the RAW image is restored using the image restoration network to obtain the restored image corresponding to the RAW image.

2. The method according to claim 1, characterized in that, The step of performing image restoration on the RAW image based on the guiding features and according to the image restoration network includes: Based on the guiding features, at least some of the features extracted from the RAW image are processed according to the image restoration network.

3. The method according to claim 1 or 2, characterized in that, The number of target layers in the image restoration network is M, and the guiding features include guiding feature maps corresponding to each of the m target layers in the M target layers, where M is a positive integer greater than or equal to 2, and m is a positive integer less than or equal to M. The target layer is a convolutional layer or a pooling layer; The step of performing image restoration on the RAW image based on the guiding features and according to the image restoration network includes: During the image restoration process of the RAW image according to the image restoration network, when passing through the first target layer of m target layers, the window data of the target window of the first target layer on the input image is processed based on the guiding feature map corresponding to the first target layer; wherein, the input image is the image input to the first target layer, and the first target layer is any one of the m target layers.

4. The method according to claim 3, characterized in that, The step of processing window data during the sliding window process of the target window of the first target layer on the input image based on the guiding feature map corresponding to the first target layer includes: Determine the window data of the target window in the first target layer; wherein, the window data is the data of the location of the target window on the input image; Determine the first guiding data corresponding to the window data; wherein, the first guiding data is the data in the guiding feature map corresponding to the first target layer; The window data is processed based on the first guidance data.

5. The method according to claim 4, characterized in that, The target window has different first guidance data corresponding to window data at different positions on the input image; and / or, The input image and the guiding feature image corresponding to the first target layer are the same in planar size. The first guiding data is the data in the guiding feature image corresponding to the first target layer that is adapted to the target window position in the input image.

6. The method according to claim 3, characterized in that, The image restoration network includes an image encoding network and an image decoding network; wherein, m target layers are located in the image encoding network and / or the image decoding network.

7. The method according to claim 1, characterized in that, The step of obtaining the photoguide map corresponding to the RAW image includes: Determine the intermediate image corresponding to each of the N pixel channels in the RAW image; Based on the intermediate images corresponding to each of the N pixel channels, the photosensitive weights corresponding to each of the N pixel channels are determined; Based on the photosensitive weights corresponding to each of the N pixel channels, the corresponding pixel channels are processed to obtain the photosensitive guide map.

8. The method according to claim 1, characterized in that, Determining the spatial guiding map corresponding to the RAW image includes: Based on the RAW image, a position encoding map corresponding to each of the N pixel channels is obtained, and the position encoding map indicates the position number of the corresponding pixel channel; Determine the position weights corresponding to each of the N pixel channels; Based on the position weights of the N pixel channels, the corresponding position encoding map is corrected to obtain the spatial guidance map.

9. The method according to claim 1, characterized in that, The photosensitive guide image is an image captured by a brightness sensor of the target scene.

10. An image restoration device, characterized in that, include: An image acquisition module is used to acquire RAW images of a target scene captured by an image sensor. The image sensor acquires RAW images through a pre-positioned color filter matrix. The color filter matrix is ​​composed of target pixel units, and each target pixel unit is composed of N target pixels. Target pixels at the same position in each target pixel unit of the color filter matrix form a pixel channel, resulting in a total of N pixel channels. The first guide map determination module is used to determine the photosensitive guide map corresponding to the RAW image, indicating the brightness information of the RAW image; The second guide map determination module is used to determine the spatial guide map corresponding to the RAW image; wherein, the spatial guide map indicates the positional relationship of the channels of the N target pixels; The feature extraction module is used to extract features based on the photosensitive guidance map and the spatial guidance map to determine the guidance features of the image restoration network; The image restoration module is used to perform image restoration on the RAW image based on the guiding features and according to the image restoration network to obtain the restored image corresponding to the RAW image.

11. The apparatus according to claim 10, characterized in that, The image restoration module is used to process at least some of the features extracted from the RAW image based on the guiding features and according to the image restoration network.

12. The apparatus according to claim 10, characterized in that, The number of target layers in the image restoration network is M, and the guiding features include guiding feature maps corresponding to each of the m target layers out of the M target layers, where M is a positive integer greater than or equal to 2 and m is a positive integer less than or equal to M; the target layers are convolutional layers or pooling layers. The image restoration module is used to process the window data of the target window of the first target layer on the input image based on the guiding feature map corresponding to the first target layer when the image is restored to the RAW image according to the image restoration network. The input image is the image input to the first target layer, and the first target layer is any one of the m target layers.

13. The apparatus according to claim 10, characterized in that, The second guide map determination module is used to obtain position encoding maps corresponding to each of the N pixel channels based on the RAW image. The position encoding maps indicate the position number of the corresponding pixel channel and the arrangement of the corresponding pixel channel in the target pixels of the color filter matrix. Based on the position encoding maps corresponding to each of the N pixel channels, the module obtains the position weights corresponding to each of the N pixel channels. Based on the position weights of the N pixel channels, the corresponding position encoding map is corrected to obtain the spatial guidance map.

14. An image restoration apparatus, characterized in that, include: At least one memory for storing programs; At least one processor is configured to execute a program stored in the memory, wherein when the program stored in the memory is executed, the processor is configured to perform the method as described in any one of claims 1-9.

15. An image restoration apparatus, characterized in that, The device executes computer program instructions to perform the method as described in any one of claims 1-9.

16. A computer storage medium, characterized in that, The computer storage medium stores instructions that, when executed on the computer, cause the computer to perform the method as described in any one of claims 1-9.

17. A computer program product containing instructions, characterized in that, When the instructions are executed on a computer, the computer causes the computer to perform the method as described in any one of claims 1-9.

Citation Information

Patent Citations

  • Image processing method and device

    CN112529775A

  • Restoration for Video Coding with Self-guided Filtering and Subspace Projection

    US20180131968A1