Image processing method, model training method, electronic equipment and storage medium

By building a lightweight neural network model, using frame accumulation and feature extraction technology to generate grid sampling features with different resolutions, the problem of high image noise in Monte Carlo path tracking and rendering of mobile electronic devices is solved, and efficient image denoising processing on mobile terminals is realized.

CN120259504AActive Publication Date: 2025-07-04HONOR DEVICE CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202311837641.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-12-27
Publication Date
2025-07-04
Estimated Expiration
2043-12-27

AI Technical Summary

Technical Problem

When mobile electronic devices use Monte Carlo paths to track rendered images, due to hardware computing power, it is difficult to take into account both rendering speed and effect, resulting in high image noise. The existing neural network methods are difficult to deploy on mobile and take up too much resources.

Method used

A lightweight neural network model is built, and grid sampling features with different resolutions are generated through frame accumulation and feature extraction, and image denoising is performed by combining weight parameters to adapt to the computing power of mobile electronic devices and realize image denoising.

Benefits of technology

It effectively reduces image noise and improves rendering effect on mobile devices. It has a small model scale and low resource usage, which is suitable for the computing power limitation of mobile devices.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120259504A_ABST
    Figure CN120259504A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides an image processing method, a model training method, electronic equipment and a storage medium. The image processing method comprises the steps of obtaining color information corresponding to each image frame in a target image frame sequence obtained through Monte Carlo path tracking rendering; determining target color information corresponding to the image frame i according to the first color information of the image frame (i-1) and the second color information of the image frame i; determining a plurality of color features and a plurality of weight parameters corresponding to the image frame i according to the target color information; generating a plurality of grid sampling features according to the first color information and the plurality of color features, the plurality of grid sampling features corresponding to different image resolution; decoding the plurality of grid sampling features by using the plurality of color features to obtain a plurality of grid decoding features, the plurality of grid decoding features corresponding to the same image resolution; and determining de-noised color information corresponding to the image frame i according to the plurality of grid decoding features and the plurality of weight parameters.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of image processing technologies, and in particular, to an image processing method, a model training method, an electronic device, and a storage medium. Background Art

[0002] Currently, in the fields of movie special effects, animation, games, etc., the Monte Carlo path tracing method is mostly used to render models in a three-dimensional scene to achieve a realistic rendering effect. However, for an electronic device to perform image rendering through the Monte Carlo path tracing method, it often requires a certain amount of computing power. Therefore, Monte Carlo path tracing is mostly applied to the PC side, and the usage scenarios are limited.

[0003] With the improvement of the hardware level of mobile electronic devices, mobile electronic devices can also provide corresponding computing power for Monte Carlo path tracing. However, due to the limited computing power that the hardware of mobile electronic devices can provide, when ensuring the rendering speed, mobile electronic devices often cannot balance the rendering effect, and the rendered images often have a lot of noise. Therefore, it is very necessary to provide an image denoising method that matches the computing power of mobile electronic devices for mobile electronic devices to use Monte Carlo path tracing to render realistic images. Summary of the Invention

[0004] Multiple aspects of the present application provide an image processing method, a model training method, an electronic device, and a storage medium for denoising images rendered by mobile electronic devices using Monte Carlo path tracing.

[0005] In a first aspect, an embodiment of the present application provides an image processing method, including:

[0006] Obtaining a target image frame sequence rendered through Monte Carlo path tracing, and color information respectively corresponding to each image frame in the target image frame sequence;

[0007] Determining target color information corresponding to the target image frame according to first color information corresponding to the previous image frame of the target image frame and second color information corresponding to the target image frame, where the target image frame is any image frame in the target image frame sequence except the first image frame;

[0008] Determining multiple color features corresponding to the target image frame according to the target color information, and weight parameters respectively corresponding to the multiple color features when representing the target color information;

[0009] Generating multiple grid sampling features corresponding one-to-one to the multiple color features according to the first color information and the multiple color features, where the multiple grid sampling features respectively correspond to different image resolutions;

[0010] Decode the multiple grid sampling features according to the multiple color features to obtain multiple grid decoding features corresponding to the multiple grid sampling features, where the multiple grid decoding features correspond to the same image resolution;

[0011] Determine the denoised color information corresponding to the target image frame according to the weight parameters respectively corresponding to the multiple grid decoding features and the multiple color features.

[0012] In a second aspect, an embodiment of the present application further provides an image denoising model training method, including:

[0013] Obtain a training sample set, where the training sample set includes: for the same object, a first image frame sequence and a second image frame sequence obtained by Monte Carlo path tracing rendering, and the image noise in the first image frame sequence is greater than that in the second image frame sequence;

[0014] Input the color information corresponding to each image frame in the first image frame sequence and the color information corresponding to each image frame in the second image frame sequence into an image denoising model, so that the image denoising model performs the following training operations:

[0015] Determine the target color information corresponding to a first target image frame according to the first color information corresponding to the previous image frame of the first target image frame and the second color information corresponding to the first target image frame, where the first target image frame is any image frame in the first image frame sequence except the first image frame;

[0016] Determine multiple color features corresponding to the first target image frame according to the target color information, and weight parameters respectively corresponding to the multiple color features when representing the target color information;

[0017] Generate multiple grid sampling features corresponding one-to-one to the multiple color features according to the first color information and the multiple color features, where the multiple grid sampling features respectively correspond to different image resolutions;

[0018] Decode the multiple grid sampling features according to the multiple color features to obtain multiple grid decoding features corresponding to the multiple grid sampling features, where the multiple grid decoding features correspond to the same image resolution;

[0019] Determine the denoised color information corresponding to the first target image frame according to the weight parameters respectively corresponding to the multiple grid decoding features and the multiple color features;

[0020] Determine a second target image frame in the second image frame sequence corresponding to the first target image frame;

[0021] Calculate the loss value between the denoised color information corresponding to the first target image frame and the color information corresponding to the second target image frame;

[0022] If the loss value is greater than a set threshold, adjust the model parameters of the image denoising model according to the loss value.

[0023] In a third aspect, an embodiment of the present application further provides an electronic device, including: a memory, a processor, and a communication interface; wherein, an executable code is stored on the memory, and when the executable code is executed by the processor, the image processing method described in the first aspect above, or the image denoising model training method described in the second aspect above is implemented.

[0024] In a fourth aspect, an embodiment of the present application further provides a non-transitory machine-readable storage medium, on which an executable code is stored. When the executable code is executed by a processor of an electronic device, the processor can at least implement the image processing method described in the first aspect above, or the image denoising model training method described in the second aspect above.

[0025] An embodiment of the present application provides a lightweight image denoising model for mobile electronic devices. The image denoising model denoises the target image frame sequence of Monte Carlo path tracing rendering in the following manner: First, obtain the color information corresponding to each image frame in the target image frame sequence. Then, for any image frame (i.e., the target image frame) in the target image frame sequence except the first image frame, perform frame accumulation on the color information of different frames according to the first color information corresponding to the previous image frame of the target image frame and the second color information corresponding to the target image frame, so as to determine the richer target color information after frame accumulation corresponding to the target image frame. After that, perform feature extraction on the target color information to determine multiple color features corresponding to the target image frame, and the weight parameters corresponding to the multiple color features when representing the target color information. Next, map the first color information and the multiple color features to a grid to generate multiple grid sampling features corresponding one-to-one to the multiple color features, where each grid sampling feature is obtained based on different grid sampling parameters, and the obtained multiple grid sampling features respectively correspond to different image resolutions. After obtaining multiple grid sampling features corresponding to different image resolutions, decode the multiple grid sampling features according to the multiple color features to obtain multiple grid decoding features corresponding to the multiple grid sampling features. The multiple grid decoding features have the same feature scale, that is, they correspond to the same image resolution. Finally, determine the denoised color information corresponding to the target image frame according to the multiple grid decoding features and the weight parameters corresponding to the multiple color features.

[0026] In this solution, image denoising of the target image frame is achieved by generating multiple grid sampling features with different feature scales (i.e., image resolutions). Since the image information corresponding to each grid in the grid sampling features corresponding to different image resolutions is different, when fusing the multiple grid decoding features decoded from the multiple grid sampling features, richer image semantic information can be obtained, and the denoised color information corresponding to the target image frame can be determined more accurately. BRIEF DESCRIPTION OF THE DRAWINGS

[0027] The drawings described herein are used to provide a further understanding of the present application and constitute a part of the present application. The illustrative embodiments and descriptions thereof of the present application are used to explain the present application and do not constitute an improper limitation of the present application. In the drawings:

[0028] Figure 1 is a schematic structural diagram of a mobile electronic device provided by an embodiment of the present application;

[0029] Figure 2 is a schematic diagram of a software module architecture provided by an embodiment of the present application;

[0030] Figure 3 is a schematic diagram of an application scenario of an image processing method provided by an embodiment of the present application Figure 1 ;

[0031] Figure 4 is a flowchart of an image processing method provided by an embodiment of the present application;

[0032] Figure 5 is a schematic diagram of an application scenario of an image processing method provided by an embodiment of the present application Figure 2 ;

[0033] Figure 6 is a schematic diagram of a scenario for determining target color information provided by an embodiment of the present application;

[0034] Figure 7 is a schematic diagram of a scenario for feature extraction provided by an embodiment of the present application;

[0035] Figure 8 is a schematic diagram of a scenario for grid sampling provided by an embodiment of the present application;

[0036] Figure 9 is a schematic diagram of a scenario for grid decoding provided by an embodiment of the present application;

[0037] Figure 10 is a schematic diagram of the model structure of an image denoising model provided by an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0038] To make the objectives, technical solutions and advantages of the present application clearer, the technical solutions of the present application will be clearly and completely described below in conjunction with the specific embodiments of the present application and the corresponding drawings. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present application without creative efforts shall fall within the protection scope of the present application.

[0039] It should be noted that the descriptions such as "first" and "second" in this text are used to distinguish different messages, devices, modules, etc., and do not represent the sequence, nor do they limit that "first" and "second" are of different types.

[0040] Currently, in the fields of movie special effects, animation, games, etc. that require image rendering of three-dimensional scenes, the commonly used image rendering methods include: rasterization rendering and path tracing. Among them, since path tracing can simulate the propagation of real light in a three-dimensional scene and calculate the lighting effect more precisely, its rendering effect is more realistic than rasterization rendering.

[0041] Monte Carlo path tracing rendering is a type of path tracing. Its basic idea is: starting from the viewpoint (i.e., the virtual camera), a virtual ray is emitted towards the three-dimensional scene through each pixel on the projection screen; when the ray intersects with the surface of an object in the three-dimensional scene, it is selectively randomly reflected in other directions according to the material properties of the object surface, and so on iteratively until the ray hits the light source or escapes from the three-dimensional scene; finally, using Monte Carlo integration, according to the contribution of each ray to the pixel, the color value corresponding to the pixel is determined.

[0042] In practical applications, the process of controlling the number of rays emitted by each pixel to determine the color of each pixel is called sampling. The number of emitted rays corresponding to each pixel can be identified by Samples Per Pixel (SPP for short), that is, the sampling number. For example, if one ray is emitted by each pixel, the sampling number is considered to be 1, denoted as 1 SPP. Among them, the larger the value of SPP, the more accurate the color value corresponding to the pixel determined by using Monte Carlo integration, and the less noise in the rendered image. However, the larger the value of SPP, the more rays are used for path tracing, the more computing power is required, and the greater the hardware computing burden on the electronic device.

[0043] With the improvement of the hardware level of mobile electronic devices, Monte Carlo path tracing is no longer only used on the PC side, and mobile electronic devices can also provide the corresponding computing power for Monte Carlo path tracing. However, due to the limited computing power that the hardware of mobile electronic devices can provide, in the case of consuming the same rendering time, the SPP supported by mobile electronic devices is lower than that of the PC side. For example, within the same rendering time, the PC side can easily support rendering with 8 SPP, 16 SPP, or even 24 SPP, but mobile electronic devices can only support rendering with 1 SPP. This means that the images rendered by mobile electronic devices using Monte Carlo path tracing often have a lot of noise. Therefore, denoising processing is essential for the images rendered by Monte Carlo path tracing with low SPP supported by mobile electronic devices.

[0044] Generally, denoising methods can be divided into two categories: traditional operator-based methods such as Spatiotemporal Varience Guided Filtering (SVGF for short), and deep learning methods based on neural networks such as Neural Bilateral Grid (NBG for short). Among them, operator-based methods often cannot meet complex scene changes. Therefore, deep learning methods based on neural networks are mostly used for image denoising.

[0045] However, existing deep learning methods based on neural networks are often developed for the PC side. On the one hand, there are problems such as consuming more resources and being difficult to deploy on mobile devices; on the other hand, there are problems that the hardware of mobile devices does not support. Taking NBG as an example, since the GuideNet structure in NBG is a five-layer DenseNet structure, too many branches will cause memory overflow on mobile electronic devices; in addition, the Grid Creation and Grid Slicing operations in NBG are custom cuda operators, which are not supported by the graphics cards of mobile devices, and even if the cuda operators are converted into operations supported by mobile electronic devices, it is easy to cause memory overflow after accumulation.

[0046] Therefore, in order to enable mobile electronic devices to render realistic images using Monte Carlo path tracing, it is very necessary to provide an image denoising method that matches the computing power of mobile electronic devices. For this purpose, the embodiments of this application provide an image processing method for denoising, which can be implemented by constructing a lightweight neural network model, aiming to denoise the images rendered by Monte Carlo path tracing with low SPP supported by mobile electronic devices by occupying less resources.

[0047] The image processing method provided by the embodiments of this application can be applied to mobile electronic devices, which include, but are not limited to, embedded devices with user interfaces such as mobile phones, tablet computers, laptop computers, wearable devices, and televisions.

[0048] Figure 1 It is a schematic structural diagram of a mobile electronic device provided by the embodiments of this application. Here, the mobile electronic device 100 is taken as a mobile phone as an example to introduce the mobile electronic device 100 provided by this application.

[0049] As Figure 1 shown, the mobile electronic device 100 may include a processor 110, an external memory interface 120, an internal memory 121, a universal serial bus (USB) interface 130, a charging management module 140, a power management module 141, a battery 142, an antenna 1, an antenna 2, a mobile communication module 150, a wireless communication module 160, an audio module 170, a speaker 170A, a receiver 170B, a microphone 170C, a headphone jack 170D, a sensor module 180, a button 190, a motor 191, an indicator 192, a camera 193, a display screen 194, and a subscriber identification module (SIM) card interface 195, etc. Among them, the sensor module 180 may include a pressure sensor 180A, a gyroscope sensor 180B, a barometric pressure sensor 180C, a magnetic sensor 180D, an acceleration sensor 180E, a distance sensor 180F, a proximity light sensor 180G, a fingerprint sensor 180H, a temperature sensor 180J, a touch sensor 180K, an ambient light sensor 180L, a bone conduction sensor 180M, etc.

[0050] It can be understood that the structure schematically shown in the embodiments of this application does not constitute a specific limitation on the mobile electronic device 100. In other embodiments of this application, the mobile electronic device 100 may include more or fewer components than those shown in the figure, or combine certain components, or split certain components, or have different component arrangements. The components shown in the figure may be implemented in hardware, software, or a combination of software and hardware.

[0051] The processor 110 may include one or more processing units. For example, the processor 110 may include an Application Processor (AP), a modem processor, a Graphics Processing Unit (GPU), an Image Signal Processor (ISP), a controller, a video codec, a Digital Signal Processor (DSP), and / or a baseband processor, etc. Among them, different processing units may be independent devices or integrated in one or more processors. Among them, the controller may generate operation control signals according to the instruction operation code and timing signals to complete the control of fetching and executing instructions.

[0052] A memory may also be provided in the processor 110 for storing instructions and data.

[0053] In some embodiments, the processor 110 may include one or more interfaces. The interfaces may include an Inter-integrated Circuit (I2C) interface, an Inter-integrated Circuit Sound (I2S) interface, a Pulse Code Modulation (PCM) interface, a Universal Asynchronous Receiver / Transmitter (UART) interface, a Mobile Industry Processor Interface (MIPI), a General-purpose Input / Output (GPIO) interface, a Subscriber Identity Module (SIM) interface, and / or a Universal Serial Bus (USB) interface, etc.

[0054] Among them, the I2C interface is a bidirectional synchronous serial bus, including a Serial Data Line (SDA) and a Serial Clock Line (SCL). In some embodiments, the processor 110 may include multiple groups of I2C buses. The processor 110 may be respectively coupled to the touch sensor 180K, the charger, the flash, the camera 193, etc. through different I2C bus interfaces.

[0055] The I2S interface can be used for audio communication. In some embodiments, the processor 110 may include multiple groups of I2S buses. The processor 110 can be coupled to the audio module 170 via the I2S bus to enable communication between the processor 110 and the audio module 170.

[0056] The PCM interface can also be used for audio communication to sample, quantize, and encode analog signals. In some embodiments, the audio module 170 and the wireless communication module 160 can be coupled via the PCM bus interface.

[0057] The UART interface is a general-purpose serial data bus for asynchronous communication. This bus can be a bidirectional communication bus. It converts the data to be transmitted between serial communication and parallel communication. In some embodiments, the UART interface is typically used to connect the processor 110 and the wireless communication module 160.

[0058] The MIPI interface can be used to connect the processor 110 to peripheral devices such as the display screen 194 and the camera 193. In some embodiments, the processor 110 and the camera 193 communicate via the MIPI interface to implement the shooting function of the mobile electronic device 100. The processor 110 and the display screen 194 communicate via the MIPI interface to implement the display function of the mobile electronic device 100.

[0059] The GPIO interface can be configured by software. The GPIO interface can be configured as a control signal or a data signal. In some embodiments, the GPIO interface can be used to connect the processor 110 to the camera 193, the display screen 194, the wireless communication module 160, the audio module 170, the sensor module 180, etc.

[0060] The USB interface 130 is an interface that complies with the USB standard specification. The USB interface 130 can be used to connect a charger to charge the mobile electronic device 100, and can also be used to transfer data between the mobile electronic device 100 and peripheral devices, etc.

[0061] It can be understood that the interface connection relationships between the modules illustrated in the embodiments of the present application are only illustrative and do not constitute a structural limitation on the mobile electronic device 100.

[0062] The charging management module 140 is used to receive charging input from a wireless charger or a wired charger. While the charging management module 140 charges the battery 142, it can also supply power to the electronic device through the power management module 141.

[0063] The wireless communication function of the mobile electronic device 100 can be implemented by the antenna 1, antenna 2, mobile communication module 150, wireless communication module 160, modulation and demodulation processor, and baseband processor, etc.

[0064] Antenna 1 and Antenna 2 are used for transmitting and receiving electromagnetic wave signals. Each antenna in the mobile electronic device 100 can be used to cover a single or multiple communication frequency bands. Different antennas can also be multiplexed to improve the utilization rate of the antennas.

[0065] The mobile communication module 150 can provide solutions for wireless communications including 2G / 3G / 4G / 5G, etc. applied to the mobile electronic device 100.

[0066] The wireless communication module 160 can provide solutions for wireless communications including Wireless Local Area Networks (WLAN), Bluetooth (BT), Global Navigation Satellite System (GNSS), Frequency Modulation (FM), Near Field Communication (NFC), Infrared (IR), etc. applied to the mobile electronic device 100. The wireless communication module 160 receives electromagnetic waves via Antenna 2, performs frequency modulation and filtering processing on the electromagnetic wave signals, and sends the processed signals to the processor 110. The wireless communication module 160 can also receive the signals to be transmitted from the processor 110, perform frequency modulation and amplification on them, and convert them into electromagnetic waves through Antenna 2 for radiation.

[0067] In some embodiments, Antenna 1 of the mobile electronic device 100 is coupled to the mobile communication module 150, and Antenna 2 is coupled to the wireless communication module 160, so that the mobile electronic device 100 can communicate with the network and other devices through wireless communication technologies. The wireless communication technologies can include Global System for Mobile Communications (GSM), General Packet Radio Service (GPRS), Code Division Multiple Access (CDMA), Wideband Code Division Multiple Access (WCDMA), Time-Division Code Division Multiple Access (TD-SCDMA), etc.

[0068] The mobile electronic device 100 realizes the display function through the GPU, the display screen 194, the application processor, etc. The GPU is a graphics processing unit that connects the display screen 194 and the application processor.

[0069] The display screen 194 is used to display images, videos, etc. The display screen 194 includes a display panel. The display panel can adopt a liquid crystal display (LCD for short), an organic light-emitting diode (OLED for short), an active matrix organic light-emitting diode or an active-matrix organic light-emitting diode (AMOLED for short), etc. In some embodiments, the mobile electronic device 100 may include one or N display screens 194, where N is a positive integer greater than 1.

[0070] The mobile electronic device 100 can realize the shooting function through the ISP, the camera 193, the video codec, the GPU, the display screen 194, the application processor, etc. The ISP is used to process the data fed back by the camera 193.

[0071] The camera 193 is used to capture static images or videos. In some embodiments, the mobile electronic device 100 may include one or N cameras 193, where N is a positive integer greater than 1.

[0072] The external memory interface 120 can be used to connect an external memory card, such as a Micro SD card, to expand the storage capacity of the mobile electronic device 100. The external memory card communicates with the processor 110 through the external memory interface 120 to realize the data storage function.

[0073] The internal memory 121 can be used to store computer-executable program codes, and the executable program codes include instructions. The processor 110 executes various functional applications and data processing of the mobile electronic device 100 by running the instructions stored in the internal memory 121 and / or the instructions stored in the memory provided in the processor.

[0074] The mobile electronic device 100 can realize the audio function through the audio module 170, the speaker 170A, the receiver 170B, the microphone 170C, the headphone jack 170D, and the application processor, etc. Such as music playback, recording, etc.

[0075] The audio module 170 is used to convert digital audio information into an analog audio signal for output, and is also used to convert an analog audio input into a digital audio signal.

[0076] The speaker 170A, also known as the "loudspeaker", is used to convert an audio electrical signal into a sound signal. The mobile electronic device 100 can listen to music or a hands-free call through the speaker 170A.

[0077] The receiver 170B, also known as the "earpiece", is used to convert an audio electrical signal into a sound signal. When the mobile electronic device 100 answers a call or a voice message, the voice can be listened to by placing the receiver 170B close to the human ear.

[0078] The microphone 170C, also known as the "microphone" or "transmitter", is used to convert a sound signal into an electrical signal. The mobile electronic device 100 can be provided with at least one microphone 170C. In some other embodiments, the mobile electronic device 100 can be provided with two microphones 170C, which can not only collect sound signals but also implement a noise reduction function. In some other embodiments, the mobile electronic device 100 can also be provided with three, four or more microphones 170C, which can collect sound signals, reduce noise, identify the sound source, and implement a directional recording function, etc.

[0079] The headphone jack 170D is used to connect a wired headphone. The headphone jack 170D can be a USB interface 130, or a 3.5 mm Open Mobile Terminal Platform (OMTP) standard interface, or a Cellular Telecommunications Industry Association of the USA (CTIA) standard interface.

[0080] The pressure sensor 180A is used to sense a pressure signal and can convert the pressure signal into an electrical signal. In some embodiments, the pressure sensor 180A can be disposed on the display screen 194.

[0081] The gyroscope sensor 180B can be used to determine the motion posture of the mobile electronic device 100.

[0082] The barometric pressure sensor 180C is used to measure the barometric pressure. In some embodiments, the mobile electronic device 100 calculates the altitude based on the barometric pressure value measured by the barometric pressure sensor 180C to assist in positioning and navigation.

[0083] The magnetic sensor 180D includes a Hall sensor. The mobile electronic device 100 can use the magnetic sensor 180D to detect the opening and closing of a flip leather case. In some embodiments, when the mobile electronic device 100 is a flip phone, the mobile electronic device 100 can detect the opening and closing of the flip according to the magnetic sensor 180D.

[0084] The acceleration sensor 180E can detect the magnitude of the acceleration of the mobile electronic device 100 in various directions (generally three axes). It can also be used to identify the posture of the electronic device and is applied to applications such as horizontal and vertical screen switching and pedometers.

[0085] The distance sensor 180F is used to measure distance. The mobile electronic device 100 can measure distance through infrared or laser. In some embodiments, when shooting a scene, the mobile electronic device 100 can use the distance sensor 180F to measure distance to achieve fast focusing.

[0086] The proximity light sensor 180G may include, for example, a light-emitting diode (LED) and a light detector. The mobile electronic device 100 uses a photodiode to detect the infrared reflected light from nearby objects. When sufficient reflected light is detected, it can be determined that there is an object near the mobile electronic device 100, otherwise there is no object near the mobile electronic device 100.

[0087] The ambient light sensor 180L is used to sense the ambient light brightness. The mobile electronic device 100 can adaptively adjust the brightness of the display screen 194 according to the sensed ambient light brightness. The ambient light sensor 180L can also cooperate with the proximity light sensor 180G to detect whether the mobile electronic device 100 is in a pocket to prevent accidental touch.

[0088] The fingerprint sensor 180H is used to collect fingerprints. The mobile electronic device 100 can use the collected fingerprint characteristics to achieve fingerprint unlocking, access application locks, fingerprint photography, fingerprint answering of incoming calls, etc.

[0089] The temperature sensor 180J is used to detect temperature. In some embodiments, the mobile electronic device 100 executes a temperature processing strategy using the temperature detected by the temperature sensor 180J. In some embodiments, when the temperature is lower than another threshold, the mobile electronic device 100 heats the battery 142 or boosts the output voltage of the battery 142 to avoid abnormal shutdown of the mobile electronic device 100 caused by low temperature.

[0090] The touch sensor 180K, also known as a "touch control device". The touch sensor 180K can be disposed on the display screen 194, and the touch sensor 180K and the display screen 194 form a touch screen, also known as a "touch control screen". The touch sensor 180K is used to detect touch operations acting on or near it. In other embodiments, the touch sensor 180K can also be disposed on the surface of the mobile electronic device 100, at a different position from the display screen 194.

[0091] The bone conduction sensor 180M can acquire vibration signals. In some embodiments, the bone conduction sensor 180M can also contact the human pulse to receive blood pressure pulsation signals. The button 190 includes a power-on button, volume buttons, etc. The button 190 can be a mechanical button or a touch button. The mobile electronic device 100 can receive button inputs and generate key signal inputs related to the user settings and function controls of the mobile electronic device 100.

[0092] The motor 191 can generate vibration prompts. The motor 191 can be used for incoming call vibration prompts and also for touch vibration feedback. The indicator 192 can be an indicator light, which can be used to indicate the charging status, power change, and can also be used to indicate messages, missed calls, notifications, etc.

[0093] The SIM card interface 195 is used to connect the SIM card. The SIM card can be inserted into or removed from the SIM card interface 195 to achieve contact and separation from the mobile electronic device 100. The mobile electronic device 100 can support 1 or N SIM card interfaces, where N is a positive integer greater than 1.

[0094] The following image processing methods in the embodiments can all be implemented in the mobile electronic device 100 with the above hardware structure.

[0095] The software system of the above mobile electronic device 100 can adopt a layered architecture, event-driven architecture, microkernel architecture, microservices architecture, or cloud architecture. In the embodiments of the present application, the Android system with a layered architecture is taken as an example to exemplarily illustrate the software structure of the mobile electronic device 100.

[0096] The layered architecture divides the software into several layers, and each layer has a clear role and division of labor. The layers communicate through interfaces. In some embodiments, the Android system can include an application layer, an application framework layer, Android runtime, system libraries, a hardware abstraction layer (HAL), and a kernel layer. It should be noted that the embodiments of the present application take the Android system as an example. In other operating systems (such as HarmonyOS, IOS system, etc.), as long as the functions implemented by each functional module are similar to those of the embodiments of the present application, the solutions of the present application can also be implemented.

[0097] Among them, the application layer can include a series of application packages.

[0098] Figure 2 It is a schematic diagram of a software module architecture provided for the embodiments of the present application.

[0099] Such as Figure 2As shown, the application package may include applications such as camera, gallery, calendar, call, map, video application, game application, VR / AR application, etc. Of course, the application layer may also include other application packages, such as payment application, shopping application, banking application, etc., which are not limited in this application.

[0100] Video applications, game applications, and VR / AR applications can be native applications that come with the operating system, or they can also be third-party applications. Video applications can be used to play and edit videos; game applications are used to provide game services; VR / AR applications can be understood as applications that support virtual reality or augmented reality technology. In some embodiments, when an image needs to be displayed during the operation of a video application, a game application, or a VR / AR application, the image can be rendered using the image processing method provided in the embodiment of the present application.

[0101] The application framework layer provides an application programming interface (API) and a programming framework for the application programs in the application layer. The application framework layer includes some predefined functions. For example, it may include a window manager, a content provider, a resource manager, a notification manager, etc., and the embodiments of the present application do not impose any restrictions on this.

[0102] The window manager is used to manage window programs. The window manager can obtain the display screen size, determine whether there is a status bar, lock the screen, capture the screen, etc.

[0103] Content providers are used to store and retrieve data and make it accessible to applications. This data can include videos, images, audio, calls made and received, browsing history and bookmarks, phone books, etc.

[0104] The resource manager provides various resources for applications, such as localized strings, icons, images, layout files, video files, and so on.

[0105] The system library may include multiple functional modules, such as surface manager, media library, 3D graphics processing library (such as OpenGL ES), 2D graphics engine (such as SGL) and rendering engine.

[0106] The surface manager is used to manage the display subsystem and provide the fusion of 2D and 3D layers for multiple applications.

[0107] The media library supports the playback and recording of multiple common audio and video formats, as well as static image files, etc. The media library can support multiple audio and video coding formats, such as: MPEG4, H.264, MP3, AAC, AMR, JPG, PNG, etc.

[0108] The 3D graphics processing library is used to implement 3D graphics drawing, image rendering, synthesis, and layer processing, etc.

[0109] The 2D graphics engine is a drawing engine for 2D drawing.

[0110] Rendering refers to the process of generating an image from a model using software. Among them, the model is a description of a three-dimensional object (or object, 3D model, model) using a language or data structure, which includes geometric, viewpoint, texture, and lighting information, etc. The image can include a digital image or a bitmap image. The rendering engine can be understood as software that uses a rendering algorithm to generate an image from a model. In some embodiments, the rendering engine can be set in a graphics processing unit (GPU).

[0111] Android Runtime includes a core library and a virtual machine. Android runtime is responsible for the scheduling and management of the Android system. The core library contains two parts: one part is the functional functions that the Java language needs to call, and the other part is the core library of Android.

[0112] The application layer and the application framework layer run in the virtual machine. The virtual machine executes the Java files of the application layer and the application framework layer as binary files. The virtual machine is used to perform functions such as object lifecycle management, stack management, thread management, security and exception management, and garbage collection.

[0113] The kernel layer is the layer between hardware and software. The kernel layer includes at least a display driver, a sensor driver, etc. In some embodiments, the display driver is used to control the display screen to display an image; the sensor driver is used to control the operation of multiple sensors, such as controlling the operation of a pressure sensor and a touch sensor.

[0114] The hardware layer may include the hardware components of the electronic device proposed above. Exemplarily, Figure 2 shows a display screen and a GPU. Among them, the GPU may include a path tracing rendering pipeline.

[0115] The following takes a game application in the application layer as an example to introduce an application scenario of the image processing method provided in the embodiments of the present application.

[0116] Figure 3 Schematic diagram of an application scenario of an image processing method provided in the embodiments of the present application Figure 1 .

[0117] As Figure 3 shown, when the game application needs to display an image, the model data of the 3D scene can be sent to the rendering engine through the rendering framework of the application framework layer; among them, the model data of the 3D scene can include: the geometric information and material information of each model in the 3D scene, as well as the path tracing rendering pipeline identifier corresponding to each model, etc.

[0118] The rendering engine receives the model data of the 3D scene, renders the 3D scene image through Monte Carlo path tracing, and performs denoising processing on the 3D scene image rendered by Monte Carlo path tracing by using the image processing method provided in the embodiments of the present application to obtain a denoised 3D scene image. The rendering engine can upload the denoised 3D scene image to the game application through the rendering framework of the application framework layer again.

[0119] When the game application receives the rendered image, if it needs to display the denoised 3D scene image, it sends the image data of the denoised 3D scene image to the display driver through the application framework layer, and the display driver controls the display screen to display the denoised 3D scene image.

[0120] The following first specifically introduces the image processing method provided in the embodiments of the present application in conjunction with the accompanying drawings, and then describes the training process of the image denoising model for implementing the image processing method.

[0121] Figure 4 is a flowchart of an image processing method provided in the embodiments of the present application. As Figure 4 shown, the image processing method includes the following steps:

[0122] 401. Obtain a target image frame sequence obtained by rendering through Monte Carlo path tracing, and color information corresponding to each image frame in the target image frame sequence.

[0123] 402. Determine the target color information corresponding to the target image frame according to the first color information corresponding to the previous image frame of the target image frame and the second color information corresponding to the target image frame, where the target image frame is any image frame in the target image frame sequence except the first image frame.

[0124] 403. Determine multiple color features corresponding to the target image frame according to the target color information, and weight parameters corresponding to the multiple color features when representing the target color information.

[0125] 404. Generate multiple grid sampling features corresponding one-to-one to the multiple color features according to the first color information and the multiple color features, and the multiple grid sampling features respectively correspond to different image resolutions.

[0126] 405. Decode multiple grid sampling features according to multiple color features to obtain multiple grid decoding features corresponding to the multiple grid sampling features, where the multiple grid decoding features correspond to the same image resolution.

[0127] 406. Determine the denoised color information corresponding to the target image frame according to the weight parameters corresponding to the multiple grid decoding features and the multiple color features.

[0128] The image processing method provided in this embodiment is applied to a scenario where continuous multiple frames of images are rendered by low-SPP Monte Carlo path tracing. The object of image processing is the target image frame sequence obtained by Monte Carlo path tracing rendering, and the goal of image processing is to determine the correct color information corresponding to each image frame in the target image frame sequence by occupying a small amount of resources, that is, to perform denoising processing on the target image frame sequence.

[0129] Among them, the target image frame sequence includes several image frames, for example, including: image frame 1, image frame 2,..., image frame N (N is a positive integer), where the N image frames are arranged in time sequence. After the target image frame sequence is obtained by Monte Carlo path tracing rendering, each image frame in the target image frame sequence carries its corresponding color information respectively, and this color information is used for subsequent image denoising processing.

[0130] Optionally, the color information corresponding to each image frame respectively includes: the albedo, pixel depth, pixel normal, and color without albedo corresponding to each pixel in each image frame. Among them, the albedo is a physical quantity representing the ability of an object to reflect light; the pixel depth is the depth of the pixel position relative to the virtual camera during Monte Carlo path tracing rendering; the pixel depth and pixel normal play a role in calibrating the pixel position.

[0131] During the image processing process based on the color information corresponding to each image frame, since the processing process corresponding to any image frame is the same, therefore, taking the target image frame (i.e., image frame i, where the value of i can be 2, 3,..., or N) in the target frame sequence as an example, combined with Figure 5 the image processing method provided in this embodiment will be described. Figure 5 Schematic diagram of an application scenario of an image processing method provided in an embodiment of the present application Figure 2 .

[0132] Among them, for the convenience of description, the target image frame is called image frame i, the previous image frame of the target image frame is called image frame (i - 1), the color information of image frame (i - 1) is called the first color information, and the color information of image frame i is called the second color information.

[0133] To improve the image denoising effect and ensure the accuracy of the color information of the denoised image frame i, first, frame accumulation processing is performed on the color information of image frame i and image frame (i - 1), that is, the target color information corresponding to image frame i is determined according to the first color information and the second color information.

[0134] Figure 6 FIG. is a schematic diagram of a scenario for determining target color information provided by an embodiment of the present application. As Figure 6 shown, as an optional frame accumulation implementation method, the first position corresponding to the virtual camera when rendering image frame (i - 1) through Monte Carlo path tracing can be obtained first, and the second position corresponding to the virtual camera when rendering image frame i through Monte Carlo path tracing, where both the first position and the second position can be described by the position matrix corresponding to the virtual camera. Then, according to the first position and the second position, the first color information is mapped to the pixel coordinate system corresponding to the second color information to obtain the third color information corresponding to the first color information. For example, through the position matrices of the first position and the second position corresponding to the virtual camera, the position transformation matrix from the first position to the second position is determined first, and based on this position transformation matrix, the first color information is mapped to the pixel coordinate system corresponding to the second color information through the reprojection algorithm, and the mapping information of the first color information in the pixel coordinate system corresponding to the second color information is denoted as the third color information. Finally, according to the third color information and the second color information, the target color information corresponding to the target image frame is determined.

[0135] Optionally, determining the target color information corresponding to image frame i according to the third color information and the second color information includes: calculating the weighted average result of the color without albedo included in the third color information and the color without albedo included in the second color information according to a preset weight coefficient; calculating the first dot product result of the weighted average result and the light rate included in the second color information; and using the first dot product result, the albedo, pixel depth, and pixel normal included in the second color information as the target color information corresponding to image frame i.

[0136] The preset weight coefficient is used to describe the respective proportions of the color without albedo in the third color information and the color without albedo in the second color information when performing frame accumulation calculation on colors. For example, the weight coefficient corresponding to the color without albedo in the third color information is 0.2, and the weight coefficient corresponding to the color without albedo in the second color information is 0.8. Optionally, the weight coefficient can be determined according to the denoising effects corresponding to multiple graphics processes.

[0137] After determining the target color information corresponding to image frame i, further, feature extraction is to be performed on the target color information for subsequent image denoising operations based on the extracted features.

[0138] Figure 7 A schematic diagram of a feature extraction scenario provided by an embodiment of this application. As shown in Figure 7 the figure, as an optional feature extraction method, the target color information can be first normalized and channel-connected to determine the first color feature corresponding to the target color information; then, the number of channels of the first color feature is adjusted to the target number of channels through, for example, convolution operations to obtain the second color feature; finally, the second color feature is sliced by channels to determine multiple color features corresponding to image frame i ( Figure 7 the color feature 1, color feature 2,..., color feature m shown in the figure, where m is a positive integer), and the weight parameters corresponding to the multiple color features when representing the target color information ( Figure 7 the weight parameter 1, weight parameter 2,..., weight parameter m shown in the figure).

[0139] Among them, the number of multiple color features can be flexibly configured according to the computing power of the mobile electronic device. The number of multiple color features, the number of weight parameters, and the target number of channels match. For example: the sum of the number of multiple color features and the number of weight parameters is equal to the target number of channels, or the sum of the number of multiple color features and the number of weight parameters minus 1 is equal to the target number of channels, etc.

[0140] Optionally, during the feature extraction process, the number of convolution operations, the size of the convolution kernel of the convolution operation, the stride stride of the convolution operation, the padding paddign, etc. can all be flexibly configured according to the computing power of the mobile electronic device.

[0141] In order to make the multiple color features and the weight parameters corresponding to the multiple colors always positive and ensure the significance of the final color calculation, after each convolution operation, the output value is restricted between 0 and 1 through an activation function. Among them, the activation function includes but is not limited to ReLU, Leaky ReLU, PreLU, Sigmoid, tanh, SELU, etc.

[0142] After obtaining multiple color features and the weight parameters corresponding to the multiple color features respectively through feature extraction, further, based on the multiple color features and the third color feature, multiple grid sampling features of different feature scales (i.e., image resolutions) are generated to realize image denoising of image frame i at different image granularities.

[0143] Figure 8 A schematic diagram of a grid sampling scenario provided by an embodiment of this application. As shown in Figure 8 the figure, as an optional grid sampling method, the color obtained by removing the albedo in the third color information can be first channel-connected with multiple color features (color feature 1, color feature 2,..., color feature m) respectively to obtain multiple channel connection results (Figure 8 (channel connection results 1, channel connection results 2, …, channel connection results m) shown in [the figure]; then, grid sampling is respectively performed on the multiple channel connection results through, for example, convolution operation, max pooling, etc., to obtain multiple grid sampling features ( Figure 8 grid sampling feature 1, grid sampling feature 2, …, grid sampling feature m) shown in [the figure].

[0144] Among them, when performing grid sampling, the grid sampling parameters corresponding to each channel connection result are different. Optionally, the grid sampling parameters include: downsampling coefficient and channel adjustment coefficient.

[0145] In the above process of grid sampling, grid sampling is respectively performed on different color features through different downsampling coefficients and channel adjustment coefficients, so as to filter out the noise information of the image from different granularities and obtain grid sampling features corresponding to different feature sizes (i.e., different resolutions). Since in different grid sampling features, the image region (i.e., image patch) corresponding to each feature is different, when fusing multiple grid sampling features, richer image semantic information can be obtained, improving the denoising effect.

[0146] For multiple grid sampling features with different feature sizes, in order to facilitate feature fusion, it is necessary to perform decoding processing on the multiple grid sampling features.

[0147] Figure 9 A schematic diagram of a grid decoding scenario provided by an embodiment of the present application is shown in Figure 9 As shown, as an optional grid decoding method, first, through methods such as deconvolution, upsampling processing is respectively performed on the multiple grid sampling features (grid sampling feature 1, grid sampling feature 2, …, grid sampling feature m). The multiple grid sampling features after upsampling processing correspond to the same image resolution, that is, they have the same feature size. It can be understood that the upsampling coefficients corresponding to different grid sampling features are different. Then, the second dot product results of the multiple grid sampling features after upsampling processing and the corresponding multiple color features are calculated, that is, the multiple grid sampling features are decoded through the multiple color features to determine the multiple grid decoding features ( Figure 9 grid decoding feature 1, grid decoding feature 2, …, grid decoding feature m) shown in [the figure]. Among them, the multiple grid decoding features correspond to the same image resolution.

[0148] Optionally, in order to make the multiple grid decoding features always positive and ensure that the final color calculation is meaningful, after the deconvolution operation, the output value is restricted between 0 and 1 through an activation function. The activation function includes but is not limited to ReLU, LeakyReLU, PreLU, Sigmoid, tanh, SELU, etc.

[0149] Finally, fuse multiple grid decoding features. In the specific implementation process, according to the first correspondence relationship between multiple grid decoding features and multiple color features, and the second correspondence relationship between multiple color features and weight parameters, determine the weight parameters corresponding to each of the multiple grid decoding features. According to the multiple grid decoding features and the weight parameters corresponding to each of the multiple grid decoding features, determine the weighted average result of the multiple grid decoding features. According to the weighted result, determine the denoised color information corresponding to image frame i.

[0150] For ease of understanding, for example, taking the Figures 5 to 9 scenario shown in as an example, if grid decoding feature 1 corresponds to color feature 1, and color feature 1 corresponds to weight parameter 1, then it is determined that grid decoding feature 1 corresponds to weight parameter 1. By analogy, it can be determined that grid decoding feature 2 corresponds to weight parameter 2, …, grid decoding feature m corresponds to weight parameter m. Then, determine the weighted average result of the m grid decoding features = (grid decoding feature 1 ⊙ weight parameter 1) + (grid decoding feature 2 ⊙ weight parameter 2) + … + (grid decoding feature m ⊙ weight parameter m). Finally, determine that the denoised color corresponding to image frame i is the product of the weighted average result of the m grid decoding features and the albedo in the target color information.

[0151] The above is the specific description of the image processing method provided by this application embodiment. For ease of understanding, next, the image processing method provided by this application embodiment will be illustrated by substituting specific parameter values. It should be noted that the specific parameter values are only for illustrative purposes and are not limited thereto.

[0152] Assume that the obtained target image frame sequence S contains N image frames, each image frame has a low SPP (for example: 1 SPP - 8 SPP), the feature size corresponding to each image frame is h × w, and the color information corresponding to each frame image includes: the albedo a ∈ R corresponding to each pixel h×w×3 , the pixel depth d ∈ R h×w×1 , the pixel normal p ∈ R h×w×3 , and the color c ∈ R after removing the albedo h×w×3 . Then, the color information of each image frame can be represented by the feature r = (a, d, p, c), where r ∈ R h ×w×10 . Correspondingly, the target image frame sequence S can be represented as S = {(r1, t1), …, (r N , t N )}, where t is used to represent its time sequence.

[0153] 1) Determine the target color information corresponding to image frame i

[0154] When determining the target color information corresponding to image frame i through frame accumulation, first map the color information r of image frame (i - 1) i-1 to the pixel coordinate system corresponding to the color information r of image frame i i to obtain the mapping result r’ i-1 of r i-1 . Among them, r i =(a i , d i , p i , c i ), r i-1 =(a i-1 , d i-1 , p i-1 , c i-1 ), r’ i-1 =(a’ i-1 , d’ i-1 , p’ i-1 , c’ i-1 ).

[0155] Then, according to the preset weight coefficient, calculate the weighted average result Ci = Avg(c’, c) of the albedo-removed color c’ included in r’ i-1 and the albedo-removed color c included in r i-1 . i i i-1 , c i ).

[0156] After that, perform an element-by-element multiplication (i.e., dot product) of the albedo-removed color Ci and the albedo a i to obtain the color information n = a i ⊙Ci. Among them, n can also be regarded as the imaging result of image frame i after frame accumulation and before removing image noise.

[0157] Finally, use the feature r’ i =(a i , d i , p i , n) as the target color information corresponding to image frame i

[0158] 2) Feature extraction from the target color information

[0159] First, normalize (a i , d i , p i , n) and perform channel connection processing to determine the first color feature f, where f = cat(normalize(a i , d i , p i , n)), f ∈ R​​h×w×10 .

[0160] In this embodiment, it is assumed that the number of multiple color features set during the training of the image denoising model is 3, and the target number of channels matching the number of multiple color features is 5. Among them, 3 of the 5 channels respectively correspond to 3 color features, and the remaining 2 channels correspond to the weight parameters of the first and second color features among the 3 color features. It can be understood that the weight parameter of the third color feature can be obtained by subtracting the weight parameters of the first and second color features from 1.

[0161] Assume that convolution operations are performed through two 5×5 convolutions to adjust the number of channels 10 of the first color feature f∈R h×w×10 to the target number of channels 5 to obtain the second color feature f”∈R h×w×5 . Among them, the stride of the two 5×5 convolutions is 1, the padding is 2, the output channel k1 of the first convolution is 20, the output channel k2 of the second convolution is 5, the activation function after the first convolution is Leaky Relu, and the activation function after the second convolution is Sigmoid. Then, the shape change of the first color feature f during the two convolution processes is as follows:

[0162] f∈R h×w×10 →f’∈R h×w×20 →f”∈R h×w×5

[0163] After that, feature slicing processing is performed on the second color feature f”∈R h×w×5 by channel, that is:

[0164] slice(f”)=(p1,p2,p3,w1,w2)

[0165] where (p1,p2,p3,w1,w2)∈R h×w×1 , p1, p2, and p3 are 3 color features corresponding to the image frame i, w1 is the weight parameter corresponding to p1 when representing the target color information (a i ,d i ,p i ,n), and w2 is the weight parameter corresponding to p2 when representing the target color information (a i ,d i ,p i ,n). The weight parameter corresponding to p3 when representing the target color information (a i ,d i ,p i ,n) is (1 - w1 - w2).

[0166] 3) Grid sampling

[0167] Suppose that for the above three color features, the three grid sampling parameters set are as follows: p1 corresponds to a downsampling factor ss = 4 and a channel adjustment factor sr = 4; p2 corresponds to a downsampling factor ss = 8 and a channel adjustment factor sr = 8; p3 corresponds to a downsampling factor ss = 16 and a channel adjustment factor sr = 16.

[0168] During grid sampling, the processing procedures corresponding to p1, p2, and p3 are the same. Taking p1 as an example, the generation process of multiple grid sampling features corresponding to multiple color features is described.

[0169] First, remove the color c’ of the albedo from the third color information r’ i-1 =(a’ i-1 , d’ i-1 , p’ i-1 , c’ i-1 ) and perform channel connection with p1. i-1 Then, pass the channel connection result through a coordinate convolution of, for example, 3×3, and the output channel is 256 / sr, that is, 256 / 4, to obtain the feature f”’ = conv(cat(c’

[0170] , p1, x, y)), where x and y respectively represent the matrices formed by the x and y coordinates in the coordinate convolution. i-1 , p1, x, y)), where x and y respectively represent the matrices formed by the x and y coordinates in the coordinate convolution.

[0171] Finally, obtain the grid sampling feature g1 corresponding to p1 through max pooling of the feature f”’.

[0172] Among them, That is

[0173] By analogy, the grid sampling feature g2 corresponding to p2 can be obtained. And the grid sampling feature g3 corresponding to p3. It can be seen that the above three grid features g1, g2, and g3 correspond to different feature sizes and channels.

[0174] 4) Grid decoding

[0175] In the grid decoding stage, use p1 to decode g1, p2 to decode g2, and p3 to decode g3. Since the decoding process of any grid sampling feature is the same, taking the decoding of g1 with p1 as an example, the process of obtaining multiple grid decoding features by decoding multiple grid sampling features is described.

[0176] First, perform coordinate deconvolution on g1 with a convolution kernel size of ss×ss (i.e., 4×4), a stride of ss = 4, and an output channel of 1, and then pass it through the activation function sigmoid to output the feature f''. The determination process of f'' can be expressed as:

[0177]

[0178] where f'' ∈ R h×w×1 , and x and y respectively represent the matrices formed by the x and y coordinates in the coordinate convolution.

[0179] After that, perform element-wise multiplication (i.e., dot product) of p1 and f'' to obtain the grid decoding feature df1 = f'' ⊙ p1, where df1 ∈ R h×w×1 .

[0180] By analogy, the decoding feature df2 corresponding to p2 decoding g2 can be obtained, where df2 ∈ R h×w×1 ; the decoding feature df3 corresponding to p3 decoding g3, where df3 ∈ R h×w×1 .

[0181] 5) Fuse the grid decoding features

[0182] First, according to the corresponding relationships between df1, df2, and df3 and p1, p2, and p3, and the relationships between p1, p2, and p3 and w1, w2, and (1 - w1 - w2), determine that the weight parameter corresponding to df1 is w1, the weight parameter corresponding to df2 is w2, and the weight parameter corresponding to df3 is (1 - w1 - w2).

[0183] Then, perform weighted average calculation on df1, df2, and df3 to generate the color information cn after denoising and removing the albedo. The calculation process can be expressed as:

[0184] cn = (df1 ⊙ w1) + (df2 ⊙ w2) + (df3 ⊙ (1 - w1 - w2))

[0185] where ⊙ represents dot product and cn ∈ R h×w×1 .

[0186] Finally, multiply cn by the color information r' of image frame i i = (a i , d i , p i , n) to obtain the color res = cn × a after denoising and including the albedo i , completing the denoising of image frame i. i

[0187] ​In summary, in the image processing method provided in this embodiment, by generating multiple grid sampling features with different feature scales (i.e., image resolutions), image denoising of the target image frame is achieved at different image granularities; since the image information corresponding to each grid in the grid sampling features corresponding to different image resolutions is different, therefore, when fusing the multiple grid decoding features decoded from the multiple grid sampling features, richer image semantic information can be obtained, and the denoised color information corresponding to the target image frame can be determined more accurately. In addition, since the amount of computation involved in the image processing method provided in the embodiment is small, even a mobile electronic device can provide sufficient computing power to execute the image processing method.

[0188] It should be noted that when introducing the image processing method in the foregoing embodiment, multiple image frames are taken as examples for illustration. In fact, the image processing method provided in this embodiment can also be used for denoising a single-frame image, and only needs to replace the color of the albedo removal corresponding to the frame accumulation of the image frame i and the image frame (i - 1) with the color of the albedo removal corresponding to the image frame i. When performing single-frame image denoising, there is no longer any association between adjacent image frames.

[0189] The image processing method provided in the embodiments of the present application can also be deployed on an electronic device with high computing power such as a PC. In the actual application process, since the computing power of the PC is stronger than that of the mobile electronic device, the execution time of this image processing method will also be shortened, so that an image denoising task with higher real-time requirements can be executed.

[0190] The image processing method provided in the embodiments of the present application can be used not only for denoising the images rendered by Monte Carlo path tracing, but also for denoising the images rendered by non-Monte Carlo path tracing, and only the number of channels in the grid sampling and grid decoding processes needs to be modified (for example: changing 10 channels to 3 channels). The modified image processing method can be used in noise reduction scenarios such as image dehazing and de-raining.

[0191] The above image processing method can be implemented by an image denoising network model with a small network scale. The training process of the image denoising network model will be described below.

[0192] First, obtain a training sample set, which includes: a first image frame sequence and a second image frame sequence rendered by Monte Carlo path tracing for the same object, and the image noise in the first image frame sequence is greater than that in the second image frame sequence.

[0193] Specifically, the first image frame sequence and the second image frame sequence can be: two image frame sequences obtained by Monte Carlo path tracing rendering with different SPPs for the same target scene. Among them, the SPP corresponding to the second image frame sequence is much larger than the SPP corresponding to the first image frame sequence. For example, the SPP corresponding to the second image frame sequence = 4096, and the SPP corresponding to the first image frame sequence = 1. The second image frame sequence can be regarded as a better result after image denoising, that is, the label in the training process of the image denoising model.

[0194] Then, input the color information corresponding to each image frame in the first image frame sequence and the color information corresponding to each image frame in the second image frame sequence into the image denoising model, so that the image denoising model performs a training operation.

[0195] For ease of understanding, it will be described in combination with the network structure of the image denoising network. Figure 10 It is a schematic diagram of the model structure of an image denoising model provided by an embodiment of the present application. As Figure 10 shown, the image denoising model includes: a data preprocessing module, a backbone network, a grid generation network, a grid decoding network, and a fusion module.

[0196] Among them, the grid generation network is composed of multiple grid generation networks with different granularities (that is, corresponding to different grid sampling parameters) (such as Figure 10 the grid generation network 1, grid generation network 2,..., grid generation network m shown); each grid generation network corresponds to a grid decoding network, so the grid decoding network is also composed of multiple different grid decoding networks (such as Figure 10 the grid decoding network 1, grid decoding network 2,..., grid decoding network m shown).

[0197] During the training process of the image denoising network, the data preprocessing module is used to determine the target color information corresponding to the first target image frame according to the first color information corresponding to the previous image frame of the first target image frame and the second color information corresponding to the first target image frame. Among them, the first target image frame is any image frame in the first image frame sequence except the first image frame.

[0198] The backbone network is used to determine multiple color features corresponding to the first target image frame according to the target color information (such as Figure 10 the color feature 1, color feature 2,..., color feature m shown), and the weight parameters corresponding to the multiple color features when representing the target color information (such as Figure 10 the weight parameter 1, weight parameter 2,..., weight parameter m shown).

[0199] During training, the number of multiple color features can be set according to the hardware capabilities of the mobile electronic device. It can be understood that the more the number of multiple color features corresponds to, the more computing power is required when the image denoising model is used.

[0200] A grid generation network, configured to generate multiple grid sampling features corresponding one-to-one to multiple color features according to the first color information and the multiple color features, and the multiple grid sampling features respectively correspond to different image resolutions.

[0201] Specifically, each grid sampling network processes one color feature to generate the grid sampling feature corresponding to the color feature. As Figure 10 shown, the grid sampling network 1 generates the grid sampling feature 1 according to the first color information and the color feature 1; the grid sampling network 2 generates the grid sampling feature 2 according to the first color information and the color feature 2; …; the grid sampling network m generates the grid sampling feature m according to the first color information and the color feature m.

[0202] A grid decoding network, configured to decode the multiple grid sampling features according to the multiple color features to obtain multiple grid decoding features corresponding to the multiple grid sampling features, and the multiple grid decoding features correspond to the same image resolution.

[0203] Specifically, each grid decoding network decodes one grid sampling feature to generate the grid decoding feature corresponding to the grid sampling feature. As Figure 10 shown, the grid decoding network 1 decodes the grid sampling feature 1 according to the color feature 1 to generate the grid decoding feature 1; the grid decoding network 2 decodes the grid sampling feature 2 according to the color feature 2 to generate the grid decoding feature 2; …; the grid decoding network m decodes the grid sampling feature m according to the color feature m to generate the grid decoding feature m.

[0204] A fusion module, configured to determine the denoised color information corresponding to the first target image frame according to the multiple grid decoding features (such as the grid decoding feature 1, the grid decoding feature 2, …, the grid decoding feature m shown in Figure 10 and the weight parameters corresponding to the multiple color features (such as the weight parameter 1, the weight parameter 2, …, the weight parameter m shown in Figure 10 ).

[0205] Finally, determine the second target image frame corresponding to the first target image frame in the second image frame sequence; and calculate the loss value between the denoised color information corresponding to the first target image frame and the color information corresponding to the second target image frame. Specifically, calculate the loss value between the color including albedo corresponding to the denoised color information of the first target image frame and the color including albedo corresponding to the color information of the second target image frame. Optionally, the loss value can be calculated by loss functions such as Smooth L1, L2, etc.

[0206] If the loss value is greater than the set threshold, the model parameters of the image denoising model are adjusted according to the loss value; if the loss value is less than or equal to the set threshold, it is considered that the model converges and the training ends.

[0207] Optionally, during the model training process, optimizers such as Adam can be used, with its initial learning rate set to lr = 1e-4 and eps = 1e-8, and a total of 100 epochs are trained, and the learning rate decays by 0.75 every 25 epochs. Optimizers such as SGD and RMSprop can also be used, and their corresponding decay strategies and epochs can also be flexibly set.

[0208] This embodiment introduces the training process of the image denoising model from the perspective of the model structure. Since the data processing process during the model training process is similar to that during the model usage process, therefore, the data processing process during the model training process will not be elaborated in this embodiment, and its specific implementation process can refer to the foregoing embodiment of the image processing method.

[0209] This embodiment provides an image denoising model and a training method for the image denoising model. Since the network scale of this image denoising model is small, the training speed is relatively fast, and only a small amount of computing power is required to deploy this image denoising model, so that it can be deployed on a mobile electronic device to perform image denoising on images rendered by Monte Carlo path tracing.

[0210] Optionally, after the image denoising model is trained, it can be deployed on a mobile electronic device through tools such as the Snapdragon Neural Processing Engine (abbreviated as SNPE), Quantum Neural Networks (abbreviated as QNN), TensorFlow Lite (abbreviated as TFLite), and coreML.

[0211] Some embodiments of this application provide an electronic device, which may include: a memory, a processor, and a communication interface. An executable code is stored on the memory. When the executable code is executed by the processor, the processor can at least implement the image processing method or the image denoising model training method provided in the foregoing embodiments. The structure of the electronic device can refer to Figure 1 the structure of the mobile electronic device 100 shown.

[0212] In addition, embodiments of this application also provide a non-transitory machine-readable storage medium, on which an executable code is stored. When the executable code is executed by the processor of the electronic device, the processor can at least implement the image processing method provided in the foregoing embodiments.

[0213] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memory, CD-ROM, optical memory, etc.) containing computer-usable program code.

[0214] The present application is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to the embodiments of the present application. It should be understood that each flow and / or block in the flowchart and / or block diagram, as well as the combination of flows and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate a device for implementing the specified functions in Figure 1 one or more flows and / or blocks Figure 1 one or more blocks.

[0215] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer-readable memory generate a manufactured article including an instruction device, and the instruction device implements the specified functions in Figure 1 one or more flows and / or blocks Figure 1 one or more blocks.

[0216] These computer program instructions can also be loaded onto a computer or other programmable data processing device, so that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process, and thus the instructions executed on the computer or other programmable device provide steps for implementing the specified functions in Figure 1 one or more flows and / or blocks Figure 1 one or more blocks.

[0217] In a typical configuration, a computing device includes one or more processors (CPUs), an input / output interface, a network interface, and a memory.

[0218] The memory may include non-permanent memory in the form of computer-readable media, such as random access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash RAM. The memory is an example of computer-readable media.

[0219] Computer-readable media includes both permanent and non-permanent, removable and non-removable media and can store information by any method or technology. The information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassettes, magnetic tape disk storage or other magnetic storage devices, or any other non-transitory medium that can be used to store information that can be accessed by a computing device. As defined herein, computer-readable media does not include transitory media such as modulated data signals and carrier waves.

[0220] It should also be noted that the term "comprising", "including" or any other variation thereof is intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements does not include only those elements but also other elements not expressly listed, or elements that are inherent to such process, method, article, or apparatus. Without further limitation, an element defined by the statement "comprising an..." does not exclude the presence of additional identical elements in the process, method, article, or apparatus that comprises the element.

[0221] The above description is only for the embodiments of the present application and is not intended to limit the present application. For those skilled in the art, various changes and modifications can be made to the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present application shall be included within the scope of the claims of the present application.

Claims

1. An image processing method, characterized in that, Including: Obtaining a target image frame sequence rendered by Monte Carlo path tracing, and color information corresponding to each image frame in the target image frame sequence; Determining target color information corresponding to a target image frame according to first color information corresponding to a previous image frame of the target image frame and second color information corresponding to the target image frame, where the target image frame is any image frame other than the first image frame in the target image frame sequence; Determining a plurality of color features corresponding to the target image frame according to the target color information, and weight parameters corresponding to the plurality of color features when representing the target color information; Generating a plurality of grid sampling features corresponding one-to-one to the plurality of color features according to the first color information and the plurality of color features, where the plurality of grid sampling features respectively correspond to different image resolutions; Decoding the plurality of grid sampling features according to the plurality of color features to obtain a plurality of grid decoding features corresponding to the plurality of grid sampling features, where the plurality of grid decoding features correspond to the same image resolution; Determining denoised color information corresponding to the target image frame according to the plurality of grid decoding features and weight parameters corresponding to the plurality of color features respectively.

2. The method according to claim 1, characterized in that, The color information corresponding to each image frame respectively includes: albedo, pixel depth, pixel normal, and color without albedo corresponding to each pixel in each image frame; where the pixel depth is the depth of the pixel position relative to the virtual camera during the Monte Carlo path tracing rendering.

3. The method according to claim 2, wherein The determining the target color information corresponding to the target image frame according to the first color information corresponding to the previous image frame of the target image frame and the second color information corresponding to the target image frame includes: Obtaining a first position of the virtual camera corresponding to rendering the previous image frame of the target image frame by the Monte Carlo path tracing, and a second position of the virtual camera corresponding to rendering the target image frame by the Monte Carlo path tracing; Mapping the first color information to a pixel coordinate system corresponding to the second color information according to the first position and the second position to obtain third color information corresponding to the first color information; Determining the target color information corresponding to the target image frame according to the third color information and the second color information.

4. The method according to claim 3, wherein The determining the target color information corresponding to the target image frame according to the third color information and the second color information includes: Calculating a weighted average result of the color without albedo included in the third color information and the color without albedo included in the second color information according to a preset weight coefficient; Calculating a first dot product result of the weighted average result and the light rate included in the second color information; Taking the first dot product result and the albedo, pixel depth, and pixel normal included in the second color information as the target color information corresponding to the target image frame.

5. The method according to claim 1, characterized in that, Determining a plurality of color features corresponding to the target image frame according to the target color information, and weight parameters corresponding to the plurality of color features when representing the target color information, includes: Extracting a first color feature corresponding to the target color information; Adjusting the number of channels of the first color feature to a target number of channels to obtain a second color feature; Performing feature slicing on the second color feature according to channels to determine a plurality of color features corresponding to the target image frame, and weight parameters corresponding to the plurality of color features when representing the target color information; the number of the plurality of color features, the number of the weight parameters, and the target number of channels match.

6. The method according to claim 3, characterized in that Generating a plurality of grid sampling features corresponding one-to-one to the plurality of color features according to the first color information and the plurality of color features, includes: Performing channel connection on the color without albedo included in the third color information and the plurality of color features respectively to obtain a plurality of channel connection results; Performing grid sampling on the plurality of channel connection results respectively to obtain a plurality of grid sampling features; wherein, when performing the grid sampling, the grid sampling parameters corresponding to each channel connection result are different.

7. The method according to claim 6, characterized in that, The grid sampling parameters include a downsampling coefficient and a channel adjustment coefficient.

8. The method according to claim 1, wherein Decoding the plurality of grid sampling features according to the plurality of color features to obtain a plurality of grid decoding features corresponding to the plurality of grid sampling features, includes: Performing upsampling processing on the plurality of grid sampling features respectively, and the plurality of upsampled grid sampling features correspond to the same image resolution; Calculating second dot product results of the plurality of upsampled grid sampling features and the corresponding plurality of color features respectively to determine a plurality of grid decoding features corresponding to the plurality of upsampled grid sampling features.

9. The method according to claim 1, wherein Determining the denoised color information corresponding to the target image frame according to the plurality of grid decoding features and the weight parameters corresponding to the plurality of color features respectively, includes: Determining weight parameters corresponding to the plurality of grid decoding features respectively according to a first correspondence between the plurality of grid decoding features and the plurality of color features, and a second correspondence between the plurality of color features and the weight parameters; Determining a weighted average result of the plurality of grid decoding features according to the plurality of grid decoding features and the weight parameters corresponding to the plurality of grid decoding features respectively; Determining the denoised color information corresponding to the target image frame according to the weighted result.

10. A method for training an image denoising model, characterized in that, Includes: Obtaining a training sample set, where the training sample set includes: a first image frame sequence and a second image frame sequence rendered by Monte Carlo path tracing for the same object, and the image noise in the first image frame sequence is greater than that in the second image frame sequence; Inputting the color information corresponding to each image frame in the first image frame sequence and the color information corresponding to each image frame in the second image frame sequence into an image denoising model, so that the image denoising model performs the following training operations: Determine the target color information corresponding to the first target image frame according to the first color information corresponding to the previous image frame of the first target image frame and the second color information corresponding to the first target image frame, where the first target image frame is any image frame in the first image frame sequence except the first image frame; Determine multiple color features corresponding to the first target image frame and weight parameters corresponding to the multiple color features when representing the target color information according to the target color information; Generate multiple grid sampling features corresponding one-to-one to the multiple color features according to the first color information and the multiple color features, where the multiple grid sampling features respectively correspond to different image resolutions; Decode the multiple grid sampling features according to the multiple color features to obtain multiple grid decoding features corresponding to the multiple grid sampling features, where the multiple grid decoding features correspond to the same image resolution; Determine the denoised color information corresponding to the first target image frame according to the multiple grid decoding features and the weight parameters corresponding to the multiple color features; Determine the second target image frame corresponding to the first target image frame in the second image frame sequence; Calculate the loss value between the denoised color information corresponding to the first target image frame and the color information corresponding to the second target image frame; If the loss value is greater than a set threshold, adjust the model parameters of the image denoising model according to the loss value.

11. An electronic device, characterized in that, Includes: A memory, a processor, and a communication interface; wherein, executable code is stored on the memory, and when the executable code is executed by the processor, the processor executes the image processing method according to any one of claims 1 to 9, or the image denoising model training method according to claim 10.

12. A non-transitory machine-readable storage medium, characterized in that, Executable code is stored on the non-transitory machine-readable storage medium, and when the executable code is executed by the processor of the electronic device, the processor executes the image processing method according to any one of claims 1 to 9, or the image denoising model training method according to claim 10.

Citation Information

Patent Citations

  • Monte Carlo rendering graph denoising method based on generative adversarial network

    CN114331895A

  • De-noising images drawn using monte carlo drawing

    CN114549374A

  • Medical image processing method and system and storage medium

    CN115965551A

  • Image noise reduction method, chip and electronic equipment

    CN116843575A

  • Rendering Method and Apparatus, and Device

    US20230230311A1