Method for removing crosstalk artifacts and related equipment

By using preset neural network models in mobile phone cameras to detect and remove crosstalk artifact areas in RAW images, the crosstalk artifact problem caused by multi-Bayer array technology is solved, and the maximum protection and improvement of image reality and clarity is achieved.

CN120075625AActive Publication Date: 2025-05-30HONOR DEVICE CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202311582092.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-11-23
Publication Date
2025-05-30
Estimated Expiration
2043-11-23

AI Technical Summary

Technical Problem

Multi-Bayer array technology causes the spacing between adjacent pixels to decrease in mobile phone cameras, increasing crosstalk problems, resulting in image signal processors being prone to crosstalk artifacts when converting RAW images to RGB images, affecting the authenticity of the image and the user's shooting experience.

Method used

The crosstalk artifact area in the RAW image acquired by the image sensor is detected and removed by the preset neural network model, and only the crosstalk artifact area is processed, and the clarity of the crosstalk artifact area is retained, thereby obtaining an RGB image after removing the crosstalk artifact.

Benefits of technology

Effectively remove crosstalk artifacts in the image, protect and restore the realism and clarity of the image to the greatest extent, and enhance the user's shooting experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120075625A_ABST
    Figure CN120075625A_ABST
Patent Text Reader

Abstract

The invention discloses a method for removing crosstalk artifacts and related equipment, which are applied to electronic equipment, the electronic equipment comprises a display screen and a camera, the camera comprises an image sensor, and pixels of the image sensor are arranged in a multi-Bayer array mode; the method comprises the following steps: receiving a first operation for starting a camera application; in response to the first operation, the display screen displays a first interface; receiving a second operation acting on the shooting control; in response to the second operation, acquiring a first image based on the first picture; detecting whether the first image comprises crosstalk artifacts or not; if the first image comprises a crosstalk artifact, correcting the first image based on a preset neural network model to obtain a second image; and based on the second image, performing conversion to obtain a third image. By adopting the embodiment of the invention, the crosstalk artifacts in the image can be removed to improve the shooting experience feeling of a user.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of image processing, and in particular, to a method and related device for removing crosstalk artifacts. Background Art

[0002] The multi-Bayer array is a technology applied to image sensors, such as Complementary Metal Oxide Semiconductor (CMOS) sensors, and is usually widely used in the cameras of electronic devices, such as mobile phone cameras. Due to the size limitation of mobile phone devices on the sensor area, it is difficult to accommodate a larger-sized sensor. Therefore, mobile phone cameras adopt the multi-Bayer array technology, introducing multiple color filters of the same color on the single-color area of each pixel in the traditional way, so as to increase the number of perceivable pixels, improve the sensitivity of the mobile phone camera, and enhance the shooting performance.

[0003] However, since the multi-Bayer array places multiple pixels of the same color on the size of the original pixel, the distance between adjacent pixels becomes smaller, which exacerbates the crosstalk problem between pixels, resulting in crosstalk artifacts easily occurring when the Image Signal Processor (ISP) converts the RAW image into a Red Green Blue (RGB) image, such as grid-like crosstalk artifacts.

[0004] Currently, the industry mainly solves this problem from two levels: hardware and software. At the hardware level, usually more advanced CMOS manufacturing processes are adopted to reduce cross-interference; at the software level, cross-interference algorithms are used to calibrate the differences between pixels under gold standard conditions, which can compensate for the differences between pixels to a certain extent. However, at the hardware level, better processes usually come with higher costs; the software-level solution can only alleviate the problem, and in some extreme cases, such as the existence of point light sources with high contrast to the background (such as the sun or lights at night), there may still be a phenomenon that crosstalk artifacts appear in the image due to crosstalk between pixels of the same color, thus reducing the authenticity of the image and affecting the user's shooting experience. Summary of the Invention

[0005] Embodiments of this application provide a method and related device for removing crosstalk artifacts, which can remove crosstalk artifacts in an image to enhance the user's shooting experience.

[0006] In a first aspect, an embodiment of the present application provides a method for removing crosstalk artifacts, which is applied to an electronic device. The electronic device includes a display screen and a camera, and the camera includes an image sensor. The pixels of the image sensor are arranged in a multi-Bayer array. The method may include: receiving a first operation to start a camera application; in response to the first operation, the display screen displays a first interface; the first interface includes a viewfinder and a shooting control. The viewfinder is used to display a first picture in real time, and the first picture is an RGB image of the current shooting scene converted from the RAW image collected by the image sensor in real time; receiving a second operation on the shooting control; in response to the second operation, obtaining a first image; the first image is the RAW image corresponding to the first picture collected by the image sensor after receiving the second operation; detecting whether the first image includes crosstalk artifacts; if the first image includes crosstalk artifacts, then based on a preset neural network model, correcting the first image to obtain a second image; the neural network model is used to remove the crosstalk artifacts in the area with crosstalk artifacts in the first image; the second image is the RAW image after removing the crosstalk artifacts; based on the second image, converting to obtain a third image; the third image is the RGB image after removing the crosstalk artifacts.

[0007] In an embodiment of the present application, when an electronic device detects that a raw image (i.e., the first image) obtained based on an image sensor and not yet processed includes crosstalk artifacts, the electronic device corrects the first image through a preset neural network model for removing crosstalk artifacts in the area with crosstalk artifacts in the first image, so as to obtain a raw image (i.e., the second image) with crosstalk artifacts removed. Then, based on the second raw image, an RGB image (i.e., the third image) with crosstalk artifacts removed is obtained. Specifically, the display interface (i.e., the first interface) of the electronic device after receiving the first operation to start the camera application includes a viewfinder and shooting controls. Among them, the viewfinder displays an RGB image (i.e., the first picture) of the current shooting scene obtained based on the raw image captured in real time by the image sensor (such as a CMOS sensor) in the camera. Further, when the electronic device receives the second operation acting on the shooting control, based on the first picture, the first image is obtained (i.e., the raw image corresponding to the first picture captured by the image sensor after receiving the second operation). Since the pixel arrangement of the image sensor is a multi-Bayer array, that is, multiple pixels of the same color are placed on the size of the original pixel, the distance between adjacent pixels becomes smaller, and the crosstalk problem between pixels of the same color is aggravated, resulting in crosstalk artifacts being prone to appear in the RGB image converted based on the raw image (such as the first image) by the electronic device, such as grid-like crosstalk artifacts, thus reducing the authenticity of the captured image. Therefore, in the embodiment of the present application, it is possible to first detect the raw image (i.e., the first image) captured by the image sensor after the electronic device receives the second operation, that is, to detect the first image to determine whether it includes crosstalk artifacts. Further, when it is detected that the raw image (i.e., the first image) obtained by the electronic device based on the image sensor includes crosstalk artifacts, then through the preset neural network model, only the crosstalk artifacts in the area with crosstalk artifacts in the first image are removed, and the area without crosstalk artifacts in the first image is not processed, so as to obtain a raw image (i.e., the second image) with crosstalk artifacts removed. As a result, in the RGB image (i.e., the third image) with crosstalk artifacts removed converted based on this raw image, the area with crosstalk artifacts is improved, and the area without crosstalk artifacts can continue to maintain the previous clarity, greatly restoring the authenticity and clarity of the image. Different from the drawback in the prior art that the authenticity and clarity of the image are greatly reduced due to the unified and rough removal of crosstalk artifacts from the entire image, the embodiment of the present application can maximize the protection and restoration of the authenticity and clarity of the image, improving the user's shooting experience.

[0008] In a possible implementation, the method further includes: after receiving a second operation on the shooting control, obtaining a fourth image; the fourth image being an RGB image corresponding to the first picture after receiving the second operation; the display screen displaying a second interface; the second interface including an image preview frame, and the fourth image being displayed in the image preview frame; receiving a third operation on the image preview frame; in response to the third operation, the display screen displaying a third interface; the third interface including an album display area, and the third image being displayed in the album display area.

[0009] In an embodiment of the present application, after the electronic device receives a second operation on the shooting control, based on the first picture displayed in the viewfinder in the above-mentioned first interface, while obtaining the RAW image (i.e., the first image) collected by the image sensor after receiving the second operation, the electronic device can also obtain the RGB image (i.e., the fourth image) corresponding to the first picture after receiving the second operation. Optionally, since the first picture displayed in the viewfinder after receiving the second operation can be a static photo, a dynamic photo or a video, the fourth image can be the RGB image corresponding to the latest frame of the first picture after receiving the second operation, or the RGB images corresponding to multiple frames of the first picture after receiving the second operation; correspondingly, the first image obtained based on the first picture can be the latest frame of the RAW image collected by the image sensor after receiving the second operation, or multiple frames of the RAW image collected by the image sensor after receiving the second operation. Further, the display screen of the electronic device displays a second interface including an image preview frame, where the image displayed in the image preview frame is the fourth image, that is, the RGB image corresponding to the latest frame of the first picture after the electronic device receives the second operation, or the RGB images corresponding to multiple frames of the first picture after receiving the second operation. After receiving a third operation on the image preview frame, the display screen displays a third interface including an album display area, where the image displayed in the album display area is the RGB image (i.e., the third image) after removing crosstalk artifacts. Through the embodiment of the present application, since the processing operation of removing crosstalk artifacts in the image may take time, when the user completes shooting, the user can immediately view the just-shot photo or video through the image preview frame to initially view the shooting effect, and then when the electronic device receives a third operation on the image preview frame, the user can directly view the image after removing crosstalk artifacts, thereby avoiding the situation that the shooting experience is affected because the user cannot view the shot photo in time due to waiting too long for the electronic device to process the crosstalk artifacts in the image, and improving the user's shooting experience.

[0010] In a possible implementation, the second interface further includes a preview photo saving control, and the third operation includes an operation on the preview photo saving control; alternatively, the third operation includes a viewing operation on the image preview frame.

[0011] In the embodiment of the present application, when the display screen of the electronic device displays the second interface including the image preview frame, the second interface may further include a preview photo saving control inside, so that the user can select whether to save the just-taken photo or video; correspondingly, the third operation on the image preview frame may include an operation on the preview photo saving control; alternatively, the third operation may include a viewing operation on the image preview frame, so that the electronic device can, in response to the third operation, display, through the album display frame of the third interface, the RGB image (i.e., the third image) after removing the crosstalk artifacts. Through the embodiment of the present application, the user can further select whether to save the just-taken photo or video through the preview photo saving control, making the user's photo-taking operation more intuitive and convenient, thereby enhancing the user's photo-taking experience.

[0012] In a possible implementation, the method further includes: storing the third image; receiving a fourth operation to start the gallery application; in response to the fourth operation, the display screen displays a fourth interface; the fourth interface includes a first display area for displaying thumbnails of the stored images; receiving a fifth operation on the thumbnail of the third image; in response to the fifth operation, the display screen displays a fifth interface; the fifth interface includes a second display area for displaying the third image.

[0013] In the embodiment of the present application, the electronic device can store the RGB image (i.e., the third image) after removing the crosstalk artifacts. When receiving the fourth operation to start the gallery application, in response to the fourth operation, the display screen of the electronic device displays the fourth interface including the first display area for displaying thumbnails of the stored images; further, when the electronic device receives the fifth operation on the thumbnail of the third image, in response to the fifth operation, the display screen of the electronic device displays the fifth interface including the second display area for displaying the third image. Through the embodiment of the present application, the user can view the photo or video after removing the crosstalk artifacts through the gallery application, enabling the user to view the image after removing the crosstalk artifacts more conveniently and quickly, and enhancing the user's shooting experience.

[0014] In a possible implementation manner, detecting whether the first image includes crosstalk artifacts may include: performing a Fourier transform on the first image to obtain a fifth image; the fifth image is a spectral image of the first image converted from the spatial domain to the frequency domain; determining whether there is a specific frequency intensity greater than a preset value in the fifth image; if there is a specific frequency intensity greater than the preset value, it is determined that the first image includes crosstalk artifacts.

[0015] In an embodiment of the present application, by performing a Fourier transform on the RAW image (i.e., the first image) corresponding to the first frame captured by the image sensor after the electronic device receives the second operation, a spectral image of the first image converted from the spatial domain to the frequency domain (i.e., the fifth image) is obtained. Further, by determining whether there is a specific frequency intensity greater than a preset value in the spectral image (i.e., the fifth image), it is determined whether the first image includes crosstalk artifacts. When it is detected that there is a specific frequency intensity greater than the preset value in the spectral image (i.e., the fifth image), it is determined that the first image includes crosstalk artifacts. Thus, through a preset neural network model, the crosstalk artifacts in the area with crosstalk artifacts in the first image can be eliminated, and the area without crosstalk artifacts is not processed, so as to retain the image clarity to the greatest extent while removing the crosstalk artifacts in the image to restore the authenticity of the image, thereby improving the user's shooting experience.

[0016] In a possible implementation manner, based on a preset neural network model, correcting the first image to obtain a second image may include: inputting the first image into the neural network model, and correcting the first image through the neural network model to obtain the second image.

[0017] In an embodiment of the present application, by inputting the RAW image (i.e., the first image) captured by the image sensor after the electronic device receives the second operation into a preset neural network model for removing crosstalk artifacts in the area with crosstalk artifacts in the first image, the electronic device can correct the first image through the neural network model to remove the crosstalk artifacts in the area with crosstalk artifacts in the first image, and obtain an RGB image without crosstalk artifacts (i.e., the second image). In an embodiment of the present application, through the neural network model, the crosstalk artifacts in the area with crosstalk artifacts in the image are specifically removed, so as to restore the image clarity and authenticity to the greatest extent, and at the same time, the efficiency of image processing can be improved, and the user's shooting experience can be enhanced.

[0018] In a possible implementation, the neural network model is trained based on a sample data set; wherein, the sample data set includes N pairs of data pairs, N being a positive integer; each data pair includes a sample image and a target image, the sample image being a RAW image with crosstalk artifacts, and the target image being a RAW image obtained by removing the crosstalk artifacts from the area with the crosstalk artifacts in the sample image based on the domain mean compensation method.

[0019] In the embodiment of the present application, the neural network model for removing the crosstalk artifacts in the area with crosstalk artifacts in the RAW image (i.e., the first image) collected by the image sensor after the electronic device receives the second operation is trained based on a sample data set. Among them, the sample data set used to train this neural network model may include N pairs of data pairs, N being a positive integer. Further, each pair of data pairs includes a sample image with crosstalk artifacts and a target image obtained by removing the crosstalk artifacts from the area with crosstalk artifacts in the sample image based on the domain mean compensation method. By training multiple times to adjust the various parameters of the neural network model, the neural network model can only remove the crosstalk artifacts in the area with crosstalk artifacts in the sample image and does not process the area without crosstalk artifacts. That is, the corrected image obtained by the neural network model based on the sample image is as close as possible to the target image corresponding to the sample image, so as to obtain a trained neural network model that can be used to remove the crosstalk artifacts in the area with crosstalk artifacts in the first image, so as to improve the image authenticity by removing the crosstalk artifacts through this neural network model, thereby enhancing the user's shooting experience.

[0020] In a possible implementation, removing the crosstalk artifacts from the area with the crosstalk artifacts in the sample image based on the domain mean compensation method may include: calculating the mean and variance of the channel values of each pixel in the pixel matrix of the same feature channel for the area with the crosstalk artifacts in the sample image; reducing the magnitude of the variance and determining the adjusted target channel value by judging whether the intensity of the specific frequency is lower than a preset value; and obtaining the target image with the crosstalk artifacts removed based on the target channel value.

[0021] In an embodiment of the present application, the crosstalk artifacts in a sample image with crosstalk artifacts are removed by a domain mean compensation method to obtain a corresponding target image after removing the crosstalk artifacts. Specifically, the mean and variance of the feature channel values of each pixel in the pixel matrix of the same feature channel in the area with crosstalk artifacts in the sample image are calculated, and then the variance is reduced, and it is determined whether the intensity of the specific frequency is lower than a preset value to adjust the channel values of each pixel in the pixel matrix, so as to obtain an adjusted feature channel value (that is, the target channel value). Further, based on the target channel value, a target image after removing the crosstalk artifacts corresponding to the sample image is obtained. Through the embodiment of the present application, the crosstalk artifacts can be removed from the area with crosstalk artifacts based on the sample image with crosstalk artifacts, so as to obtain a target image after removing the crosstalk artifacts while minimizing the loss of image sharpness. In this way, a sample data set can be made to train a neural network model to obtain an optimal neural network model for removing crosstalk artifacts, so as to improve the image authenticity and thus enhance the user's shooting experience.

[0022] In a second aspect, the present application provides a device for removing crosstalk artifacts, which is applied to an electronic device. The electronic device includes a display screen and a camera, and the camera includes an image sensor. Among them, the pixels of the image sensor are arranged in a multi-Bayer array; it may include:

[0023] A first receiving unit, configured to receive a first operation for starting a camera application;

[0024] A first display unit, configured to, in response to the first operation, display a first interface on the display screen; wherein, the first interface includes a viewfinder and shooting controls, and the viewfinder is used to display a first picture in real time, and the first picture is an RGB image of the current shooting scene converted from the RAW image collected by the image sensor in real time;

[0025] A second receiving unit, configured to receive a second operation acting on the shooting control;

[0026] A first obtaining unit, configured to, in response to the second operation, obtain a first image; the first image is the RAW image corresponding to the first picture collected by the image sensor after receiving the second operation;

[0027] A first detecting unit, configured to detect whether the first image includes crosstalk artifacts;

[0028] A first correction unit, configured to, if the first image includes crosstalk artifacts, correct the first image based on a preset neural network model to obtain a second image; wherein, the neural network model is used to remove the crosstalk artifacts in the area with crosstalk artifacts in the first image; the second image is a RAW image after removing the crosstalk artifacts.

[0029] A first conversion unit, configured to convert the second image to obtain a third image; the third image is an RGB image after removing the crosstalk artifacts.

[0030] The apparatus for removing crosstalk artifacts provided by the embodiments of the present application is applied to an electronic device. When the first receiving unit receives a first operation to start the camera application, the first display unit responds to the first operation and displays, through the display screen of the electronic device, a first interface including a viewfinder and shooting controls. Among them, the viewfinder is used to display in real time an RGB image (i.e., the first picture) of the current shooting scene converted from the RAW image collected in real time by the image sensor. Further, when the second receiving unit receives a second operation on the shooting control, the first obtaining unit responds to the second operation and obtains the RAW image (i.e., the first image) corresponding to the first picture collected by the image sensor after receiving the second operation. Still further, the first detecting unit detects whether the first image includes crosstalk artifacts. If it is detected that the first image includes crosstalk artifacts, the first correcting unit corrects the first image based on a preset neural network model to remove the crosstalk artifacts in the area with crosstalk artifacts in the first image, obtaining a RAW image (i.e., the second image) after removing the crosstalk artifacts. Then, the first converting unit converts, based on the second image, to obtain an RGB image (i.e., the third image) after removing the crosstalk artifacts. Since the pixel arrangement of the image sensor is a multi-Bayer array, that is, multiple pixels of the same color are placed on the size of the original pixel, the distance between adjacent pixels becomes smaller, and the crosstalk problem between pixels of the same color is aggravated, making it easy for crosstalk artifacts to appear in the RGB image converted by the electronic device based on the RAW image (such as the first image), such as grid-like crosstalk artifacts, thus reducing the authenticity of the captured image. Therefore, through the embodiments of the present application, it is possible to first detect the RAW image corresponding to the first picture collected by the image sensor after the electronic device receives the second operation, that is, to determine whether it includes crosstalk artifacts by detecting the first image. Further, when it is detected that the RAW image (i.e., the first image) obtained by the electronic device based on the image sensor includes crosstalk artifacts, the crosstalk artifacts in only the area with crosstalk artifacts in the first image are removed through a preset neural network model, and the area without crosstalk artifacts in the first image is not processed, obtaining a RAW image (i.e., the second image) after removing the crosstalk artifacts. As a result, in the RGB image (i.e., the third image) after removing the crosstalk artifacts converted based on this RAW image, the area with crosstalk artifacts is improved, and the area without crosstalk artifacts can continue to maintain the previous clarity, greatly restoring the authenticity and clarity of the image, thereby enhancing the user's shooting experience.

[0031] In a possible implementation manner, the apparatus further includes:

[0032] A second acquisition unit, configured to acquire a fourth image after receiving a second operation on the photographing control; the fourth image is an RGB image corresponding to the first picture after receiving the second operation.

[0033] A second display unit, configured to display a second interface on the display screen; the second interface includes an image preview frame, and the fourth image is displayed in the image preview frame.

[0034] A third receiving unit, configured to receive a third operation on the image preview frame.

[0035] A third display unit, configured to, in response to the third operation, display a third interface on the display screen; the third interface includes an album display area, and the third image is displayed in the album display area.

[0036] In a possible implementation manner, the second interface further includes a preview photo saving control, and the third operation includes an operation on the preview photo saving control; or, the third operation includes a viewing operation on the image preview frame.

[0037] In a possible implementation manner, the apparatus further includes:

[0038] A first storage unit, configured to store the third image.

[0039] A fourth receiving unit, configured to receive a fourth operation to start a gallery application.

[0040] A fourth display unit, configured to, in response to the fourth operation, display a fourth interface on the display screen; the fourth interface includes a first display area, and the first display area is used to display thumbnails of stored images.

[0041] A fifth receiving unit, configured to receive a fifth operation on the thumbnail of the third image.

[0042] A fifth display unit, configured to, in response to the fifth operation, display a fifth interface on the display screen; the fifth interface includes a second display area, and the second display area is used to display the third image.

[0043] In a possible implementation manner, the first detection unit is specifically configured to:

[0044] Perform a Fourier transform on the first image to obtain a fifth image; the fifth image is a spectrum image of the first image converted from the spatial domain to the frequency domain.

[0045] Determine whether there is a specific frequency intensity greater than a preset value in the fifth image.

[0046] If there exists a specific frequency intensity greater than the preset value, it is determined that the first image includes crosstalk artifacts.

[0047] In a possible implementation manner, the first correction unit is specifically configured to:

[0048] Input the first image into the neural network model, and correct the first image through the neural network model to obtain the second image.

[0049] In a possible implementation manner, the neural network model is trained based on a sample data set; wherein, the sample data set includes N pairs of data pairs, N is a positive integer; each data pair includes a sample image and a target image, the sample image is a RAW image with crosstalk artifacts, and the target image is a RAW image obtained by removing the crosstalk artifacts from the area with the crosstalk artifacts in the sample image based on the domain mean compensation method.

[0050] In a possible implementation manner, removing the crosstalk artifacts from the area with the crosstalk artifacts in the sample image based on the domain mean compensation method includes:

[0051] For the area with the crosstalk artifacts in the sample image, calculate the mean and variance of the channel values of each pixel in the pixel matrix of the same feature channel;

[0052] By reducing the magnitude of the variance, and determining the adjusted target channel value by judging whether the intensity of the specific frequency is lower than the preset value;

[0053] Based on the target channel value, obtain the target image with the crosstalk artifacts removed.

[0054] In a third aspect, an embodiment of the present application provides an electronic device, which may include a memory and a processor. Wherein, the memory is used to store a computer program, and the processor is used to execute the computer program stored in the memory, so that the electronic device executes the method described in any one of the second aspects above.

[0055] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, which stores a computer program, and the computer program is executed by the processor to implement the method described in any one of the first aspect or the second aspect above.

[0056] In a fifth aspect, an embodiment of the present application provides a computer program, the computer program includes instructions, and the computer program is executed by a computing device to implement the method described in any one of the first aspect or the second aspect above. Description of the Drawings

[0057] To more clearly illustrate the technical solutions in the embodiments of the present application or the background art, the following will describe the drawings required to be used in the embodiments of the present application or the background art.

[0058] Figure 1 It is a schematic diagram of the hardware structure of an electronic device in the prior art.

[0059] Figure 2 It is a schematic diagram of the software structure of the electronic device 100 provided in the embodiment of the present application.

[0060] Figure 3A It is a schematic diagram of a user interface for removing crosstalk artifacts in a photographing scene provided in the embodiment of the present application.

[0061] Figure 3B It is another schematic diagram of a user interface for removing crosstalk artifacts in a photographing scene provided in the embodiment of the present application.

[0062] Figure 3C It is a schematic diagram of a user interface for viewing an image after removing crosstalk artifacts after photographing is completed provided in the embodiment of the present application.

[0063] Figure 4A It is a flowchart example of a method for removing crosstalk artifacts provided in the embodiment of the present application.

[0064] Figure 4B It is another flowchart example of a method for removing crosstalk artifacts provided in the embodiment of the present application.

[0065] Figure 4C It is a flowchart schematic diagram for viewing an image after removing artifacts after a user completes shooting provided in the embodiment of the present application.

[0066] Figure 5 It is a schematic diagram of removing crosstalk artifacts by a domain mean compensation method provided in the embodiment of the present application.

[0067] Figure 6 It is a schematic diagram of training a neural network model based on a U-Net architecture provided in the embodiment of the present application.

[0068] Figure 7 It is a schematic diagram of the structure of a device for removing crosstalk artifacts provided in the embodiment of the present application.

[0069] Figure 8 It is another schematic diagram of the hardware structure of an electronic device provided in the embodiment of the present application. Detailed implementation manners

[0070] The following will describe the embodiments of the present application in combination with the drawings in the embodiments of the present application.

[0071] The terms used in the following embodiments of the present application are only for the purpose of describing specific embodiments and are not intended to limit the embodiments of the present application. As used in the description and appended claims of the embodiments of the present application, the singular forms "a", "an", "the", "above-mentioned", "said", and "this" are also intended to include the plural forms, unless the context clearly indicates otherwise. It should also be understood that the term "and / or" used in the embodiments of the present application refers to and includes any or all possible combinations of one or more of the listed items.

[0072] The terms "first", "second", "third", "fourth", etc. in the description and claims of the present application and the said drawings are used to distinguish different objects and not to describe a specific order. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device that includes a series of steps or units is not limited to the listed steps or units, but may optionally further include steps or units not listed, or may optionally further include other steps or units inherent to these processes, methods, products or devices.

[0073] Referring to "embodiments" herein means that specific features, structures, or characteristics described in connection with the embodiments may be included in at least one embodiment of the present application. The phrase appears in various places in the description and does not necessarily refer to the same embodiment, nor is it an independent or alternative embodiment mutually exclusive with other embodiments. Those skilled in the art will explicitly and implicitly understand that the embodiments described herein may be combined with other embodiments.

[0074] First, some terms in the present application are explained to facilitate understanding of the embodiments of the present application by those skilled in the art.

[0075] (1) Crosstalk: In the field of photography, crosstalk generally refers to unwanted signal interference on an image sensor due to the mutual influence between adjacent photosensitive elements.

[0076] (2) Crosstalk artifact: Crosstalk artifact is an interference or defect that appears in an image, usually manifested as a grid-like texture or artifact in the image. In digital photography, one common crosstalk situation is the interaction between the arrangement structure of the camera's photosensitive elements (such as sensors) and the fine patterns or textures in the object being photographed. This interaction may cause frequency aliasing, resulting in obvious grid-like artifacts.

[0077] (3) RAW image: Refers to the original image captured by an image sensor without any processing. The RAW image retains the original data of each pixel obtained from the sensor, including brightness, color, and other relevant information.

[0078] (4) Channel value: Refers to the numerical value of the color channels in a digital image. In a color image, usually the three primary color channels of red (R), green (G), and blue (B) are used to represent the color of each pixel. Each channel has a numerical value indicating the color intensity or brightness on that channel.

[0079] (5) Upsampling: Refers to increasing the sampling rate of a signal or an image, that is, increasing the number of samples. In image processing, upsampling is usually used to enlarge a low-resolution image or feature map to the original resolution for more detailed analysis or comparison with a high-resolution image. Common upsampling methods include nearest neighbor interpolation, bilinear interpolation, and transposed convolution (also known as deconvolution).

[0080] (6) Downsampling: Refers to reducing the sampling rate of a signal or an image, that is, reducing the number of samples. In image processing, downsampling is usually to reduce the size of the image, thereby reducing the computational cost, storage requirements, or simplifying subsequent processing. Common downsampling methods include average pooling (Average Pooling, AP) and max pooling (Max Pooling, MP), where the pooling operation extracts the features of the image and reduces the dimension of the image by taking the average or maximum value over an image region.

[0081] (7) Convolutional Neural Network (CNN): A type of deep learning neural network that can effectively capture and learn hierarchical feature representations, which makes it perform well in fields such as image processing and computer vision. Due to the translational invariance and parameter sharing characteristics of CNNs for image data, they can effectively process data with a spatial hierarchical structure and are mainly used for tasks that process and analyze data with a grid structure, such as image and video recognition, computer vision tasks, etc.

[0082] (8) Convolutional Layer (CL): The convolutional layer is the core part of the CNN. It uses a convolutional kernel (filter) to perform a convolution operation on the input data to extract local features in the input data. The convolution operation slides the convolutional kernel over the input data, calculates the convolution result at each position, and generates an output feature map.

[0083] (9) Pooling Layer (PL): The pooling layer is used to reduce the spatial dimension of the feature map, reduce the computational complexity, and improve the robustness of the model. Max pooling and average pooling are common pooling operations used to retain the most significant features.

[0084] (10) Loss Function (LF): Defines the difference between the output of the model for a given input and the actual label.

[0085] To facilitate the understanding of the embodiments of the present application, an exemplary electronic device provided in the embodiments of the present application will be introduced first below.

[0086] Please refer to Figure 1 , Figure 1 , which is a schematic diagram of the hardware structure of an electronic device provided in the embodiments of the present application. The electronic device 100 is a smart terminal device and can be of various types. The embodiments of the present application do not limit its specific type. For example, the terminal device can be a mobile phone, and can also include a tablet computer, a desktop computer, a desktop computer with a touch-sensitive surface or touch panel, a laptop, a handheld computer, a notebook computer, a smart screen, a wearable device (such as a smart watch, a smart bracelet, etc.), an augmented reality (AR) device, a virtual reality (VR) device, an artificial intelligence (AI) device, a car machine, a smart headset, a game console, and can also be an Internet of Things (IOT) device or a smart home device such as a smart water heater, a smart lamp, a smart air conditioner, etc.

[0087] Please refer to Figure 1 , and the following will specifically introduce each component of the electronic device 100 in conjunction with Figure 1 :

[0088] The electronic device 100 may include a processor 110, an external memory interface 120, an internal memory 121, a universal serial bus (USB) interface 130, a charging management module 140, a power management module 141, a battery 142, an antenna 1, an antenna 2, a mobile communication module 150, a wireless communication module 160, an audio module 170, a speaker 170A, a receiver 170B, a microphone 170C, a headphone jack 170D, a sensor module 180, a button 190, a motor 191, an indicator 192, a camera 193, a display screen 194, and a subscriber identification module (SIM) card interface 195, etc. Among them, the sensor module 180 may include a pressure sensor 180A, a gyroscope sensor 180B, a barometric pressure sensor 180C, a magnetic sensor 180D, an acceleration sensor 180E, a distance sensor 180F, a proximity light sensor 180G, a fingerprint sensor 180H, a temperature sensor 180J, a touch sensor 180K, an ambient light sensor 180L, a bone conduction sensor 180M, etc.

[0089] It can be understood that the structure illustrated in the embodiments of the present invention does not constitute a specific limitation on the electronic device 100. In some other embodiments of the present application, the electronic device 100 may include more or fewer components than those illustrated, or combine certain components, or split certain components, or have different component arrangements. The illustrated components may be implemented in hardware, software, or a combination of software and hardware.

[0090] The processor 110 may include one or more processing units. For example, the processor 110 may include an application processor (AP), a modem processor, a graphics processing unit (GPU), an image signal processor (ISP), a controller, a memory, a video codec, a digital signal processor (DSP), a baseband processor, and / or a neural-network processing unit (NPU), etc. Among them, different processing units may be independent devices or integrated in one or more processors.

[0091] Among them, the controller may be the nerve center and command center of the electronic device 100. The controller may generate operation control signals according to the instruction operation code and timing signal to complete the control of fetching and executing instructions.

[0092] A memory may also be provided in the processor 110 for storing instructions and data. In some embodiments, the memory in the processor 110 is a cache memory. This memory may save the instructions or data that the processor 110 has just used or recycled. If the processor 110 needs to use the instruction or data again, it can be directly called from the memory. This avoids repeated accesses, reduces the waiting time of the processor 110, and thus improves the efficiency of the system.

[0093] In some embodiments, the processor 110 may include one or more interfaces. The interfaces may include an inter-integrated circuit (I2C) interface, an inter-integrated circuit sound (I2S) interface, a pulse code modulation (PCM) interface, a universal asynchronous receiver / transmitter (UART) interface, a mobile industry processor interface (MIPI), a general-purpose input / output (GPIO) interface, a subscriber identity module (SIM) interface, and / or a universal serial bus (USB) interface, etc.

[0094] It can be understood that the interface connection relationships between the modules illustrated in the embodiments of the present invention are only illustrative descriptions and do not constitute a structural limitation on the electronic device 100. In other embodiments of the present application, the electronic device 100 may also adopt different interface connection methods in the above embodiments, or a combination of multiple interface connection methods.

[0095] The charging management module 140 is configured to receive a charging input from a charger. The charger may be a wireless charger or a wired charger. In some embodiments of wired charging, the charging management module 140 may receive the charging input of the wired charger through the USB interface 130. In some embodiments of wireless charging, the charging management module 140 may receive the wireless charging input through the wireless charging coil of the electronic device 100. While charging the battery 142, the charging management module 140 may also supply power to the electronic device through the power management module 141.

[0096] The power management module 141 is used to connect the battery 142, the charging management module 140, and the processor 110. The power management module 141 receives the inputs from the battery 142 and / or the charging management module 140 and supplies power to the processor 110, the internal memory 121, the external memory, the display screen 194, the camera 193, and the wireless communication module 160, etc. The power management module 141 may also be used to monitor parameters such as the battery capacity, the number of battery cycles, and the battery health status (leakage, impedance). In some other embodiments, the power management module 141 may also be disposed in the processor 110. In other embodiments, the power management module 141 and the charging management module 140 may also be disposed in the same device.

[0097] The wireless communication function of the electronic device 100 can be implemented by antenna 1, antenna 2, the mobile communication module 150, the wireless communication module 160, the modulation and demodulation processor, and the baseband processor, etc.

[0098] Antenna 1 and antenna 2 are used to transmit and receive electromagnetic wave signals. Each antenna in the electronic device 100 can be used to cover a single or multiple communication frequency bands. Different antennas can also be multiplexed to improve the utilization rate of the antennas. For example, antenna 1 can be multiplexed as the diversity antenna of the wireless local area network. In some other embodiments, the antenna can be used in combination with a tuning switch.

[0099] The mobile communication module 150 can provide solutions for wireless communications such as 2G / 3G / 4G / 5G applied to the electronic device 100. The mobile communication module 150 can include at least one filter, switch, power amplifier, low noise amplifier (LNA), etc. The mobile communication module 150 can receive electromagnetic waves by antenna 1, filter, amplify, etc. the received electromagnetic waves, and transmit them to the modulation and demodulation processor for demodulation. The mobile communication module 150 can also amplify the signal modulated by the modulation and demodulation processor and convert it into electromagnetic waves through antenna 1 for radiation. In some embodiments, at least some functional modules of the mobile communication module 150 can be provided in the processor 110. In some embodiments, at least some functional modules of the mobile communication module 150 and at least some modules of the processor 110 can be provided in the same device.

[0100] The modulation and demodulation processor can include a modulator and a demodulator. In some embodiments, the modulation and demodulation processor can be an independent device. In some other embodiments, the modulation and demodulation processor can be independent of the processor 110 and be provided in the same device as the mobile communication module 150 or other functional modules.

[0101] The wireless communication module 160 may provide solutions for wireless communications applied to the electronic device 100, including wireless local area networks (WLANs) (such as wireless fidelity (Wi-Fi) networks), Bluetooth (BT), global navigation satellite systems (GNSSs), frequency modulation (FM), near field communication (NFC), infrared (IR), and the like. The wireless communication module 160 may be one or more devices integrating at least one communication processing module. The wireless communication module 160 receives electromagnetic waves via the antenna 2, performs frequency modulation and filtering processing on the electromagnetic wave signals, and sends the processed signals to the processor 110. The wireless communication module 160 may also receive signals to be sent from the processor 110, perform frequency modulation and amplification on them, and convert them into electromagnetic waves through the antenna 2 for radiation.

[0102] In some embodiments, antenna 1 of electronic device 100 is coupled to mobile communication module 150, and antenna 2 is coupled to wireless communication module 160, such that electronic device 100 can communicate with a network and other devices through wireless communication technologies. The wireless communication technologies may include global system for mobile communications (GSM), general packet radio service (GPRS), code division multiple access (CDMA), wideband code division multiple access (WCDMA), time-division code division multiple access (TD-SCDMA), long term evolution (LTE), BT, GNSS, WLAN, NFC, FM, and / or IR technologies, etc. The GNSS may include global positioning system (GPS), global navigation satellite system (GLONASS), beidou navigation satellite system (BDS), quasi-zenith satellite system (QZSS), and / or satellite based augmentation systems (SBAS).

[0103] Electronic device 100 implements a display function through a GPU, display screen 194, and an application processor, etc. The GPU is a microprocessor for image processing, and is connected to display screen 194 and the application processor. The GPU is used to perform mathematical and geometric calculations for graphics rendering. Processor 110 may include one or more GPUs, which execute program instructions to generate or change display information.

[0104] The display screen 194 is used to display images, videos, etc. The display screen 194 includes a display panel. The display panel can adopt a liquid crystal display (LCD), an organic light-emitting diode (OLED), an active-matrix organic light-emitting diode (AMOLED), a flexible light-emitting diode (FLED), a Miniled, a MicroLed, a Micro-oLed, a quantum dot light-emitting diode (QLED), etc. In some embodiments, the electronic device 100 may include one or N display screens 194, where N is a positive integer greater than 1.

[0105] The electronic device 100 can implement the shooting function through an ISP, a camera 193, a video codec, a GPU, a display screen 194, an application processor, etc.

[0106] The ISP is used to process the data fed back by the camera 193. For example, when taking a photo, the shutter is opened, and light passes through the lens and is transmitted to the camera photosensitive element. The optical signal is converted into an electrical signal, and the camera photosensitive element transmits the electrical signal to the ISP for processing and converts it into an image visible to the naked eye. The ISP can also optimize the noise and brightness of the image through algorithms. The ISP can also optimize parameters such as the exposure and color temperature of the shooting scene. In some embodiments, the ISP can be set in the camera 193.

[0107] The camera 193 is used to capture static images or videos. An object generates an optical image through the lens and projects it onto the photosensitive element. The photosensitive element can be a charge-coupled device (CCD) or a complementary metal-oxide-semiconductor (CMOS) phototransistor. The photosensitive element converts the optical signal into an electrical signal, and then transmits the electrical signal to the ISP to convert it into a digital image signal. The ISP outputs the digital image signal to the DSP for processing. The DSP converts the digital image signal into an image signal in standard RGB, YUV, etc. formats. In some embodiments, the electronic device 100 may include one or N cameras 193, where N is a positive integer greater than 1.

[0108] The digital signal processor is used to process digital signals. Besides being able to process digital image signals, it can also process other digital signals. For example, when the electronic device 100 selects a frequency point, the digital signal processor is used to perform Fourier transform on the frequency point energy, etc.

[0109] The video codec is used to compress or decompress digital videos. The electronic device 100 can support one or more video codecs. In this way, the electronic device 100 can play or record videos in multiple encoding formats, such as: Moving Picture Experts Group (MPEG) 1, MPEG2, MPEG3, MPEG4, etc.

[0110] The NPU is a neural-network (NN) computing processor. By drawing on the structure of biological neural networks, such as the transmission pattern between human brain neurons, it can quickly process input information and can also continuously self-learn. Through the NPU, applications such as intelligent cognition of the electronic device 100 can be realized, such as: image recognition, face recognition, speech recognition, text understanding, etc.

[0111] The external memory interface 120 can be used to connect an external memory card, such as a Micro SD card, to expand the storage capacity of the electronic device 100. The external memory card communicates with the processor 110 through the external memory interface 120 to achieve the data storage function. For example, files such as music and videos are saved in the external memory card.

[0112] The internal memory 121 can be used to store computer-executable program code, and the executable program code includes instructions. The processor 110 executes various functional applications and data processing of the electronic device 100 by running the instructions stored in the internal memory 121. The internal memory 121 can include a program storage area and a data storage area. Among them, the program storage area can store the operating system, application programs required for at least one function (such as the sound playback function, the image playback function, etc.). The data storage area can store the data created during the use of the electronic device 100 (such as audio data, phone book, etc.). In addition, the internal memory 121 can include high-speed random access memory and can also include non-volatile memory, such as at least one disk storage device, flash memory device, universal flash storage (UFS), etc.

[0113] The electronic device 100 can implement audio functions through the audio module 170, speaker 170A, receiver 170B, microphone 170C, headphone jack 170D, and the application processor, etc. Such as music playback, recording, etc.

[0114] The audio module 170 is used to convert digital audio information into an analog audio signal for output, and is also used to convert an analog audio input into a digital audio signal. The audio module 170 can also be used for encoding and decoding audio signals. In some embodiments, the audio module 170 can be disposed in the processor 110, or some functional modules of the audio module 170 can be disposed in the processor 110.

[0115] The speaker 170A, also known as the "loudspeaker", is used to convert an audio electrical signal into a sound signal. The electronic device 100 can listen to music or hands-free calls through the speaker 170A.

[0116] The receiver 170B, also known as the "earpiece", is used to convert an audio electrical signal into a sound signal. When the electronic device 100 answers a call or a voice message, the voice can be listened to by placing the receiver 170B close to the human ear.

[0117] The microphone 170C, also known as the "microphone" or "transmitter", is used to convert a sound signal into an electrical signal. When making a call or sending a voice message, the user can speak by placing the mouth close to the microphone 170C to input the sound signal into the microphone 170C. The electronic device 100 can be provided with at least one microphone 170C. In some other embodiments, the electronic device 100 can be provided with two microphones 170C, which can not only collect sound signals but also implement a noise reduction function. In some other embodiments, the electronic device 100 can also be provided with three, four or more microphones 170C to implement functions such as collecting sound signals, noise reduction, identifying the sound source, and implementing a directional recording function.

[0118] The headphone jack 170D is used to connect a wired headphone. The headphone jack 170D can be a USB interface 130, or a 3.5 mm open mobile terminal platform (OMTP) standard interface, or a cellular telecommunications industry association of the USA (CTIA) standard interface.

[0119] The pressure sensor 180A is used to sense pressure signals and can convert pressure signals into electrical signals. In some embodiments, the pressure sensor 180A may be disposed on the display screen 194. There are many types of pressure sensors 180A, such as resistive pressure sensors, inductive pressure sensors, capacitive pressure sensors, etc. The capacitive pressure sensor may include at least two parallel plates with conductive materials. When a force acts on the pressure sensor 180A, the capacitance between the electrodes changes. The electronic device 100 determines the intensity of the pressure based on the change in capacitance. When a touch operation acts on the display screen 194, the electronic device 100 detects the intensity of the touch operation according to the pressure sensor 180A. The electronic device 100 can also calculate the position of the touch based on the detection signal of the pressure sensor 180A. In some embodiments, touch operations acting on the same touch position but with different touch operation intensities may correspond to different operation instructions. For example, when a touch operation with a touch operation intensity less than the first pressure threshold acts on the short message application icon, the instruction to view short messages is executed. When a touch operation with a touch operation intensity greater than or equal to the first pressure threshold acts on the short message application icon, the instruction to create a new short message is executed.

[0120] The gyroscope sensor 180B can be used to determine the motion posture of the electronic device 100.

[0121] The barometric pressure sensor 180C is used to measure barometric pressure.

[0122] The magnetic sensor 180D includes a Hall sensor.

[0123] The acceleration sensor 180E can detect the magnitude of the acceleration of the electronic device 100 in various directions (generally three axes). When the electronic device 100 is stationary, the magnitude and direction of gravity can be detected. It can also be used to identify the posture of the electronic device and is applied to applications such as horizontal and vertical screen switching and pedometers.

[0124] The distance sensor 180F is used to measure distance.

[0125] The proximity light sensor 180G may include, for example, a light emitting diode (LED) and a light detector, such as a photodiode.

[0126] The ambient light sensor 180L is used to sense the ambient light brightness.

[0127] The fingerprint sensor 180H is used to collect fingerprints. The electronic device 100 can use the collected fingerprint characteristics to achieve fingerprint unlocking, access application locks, fingerprint photography, fingerprint answering calls, etc.

[0128] The temperature sensor 180J is used to detect temperature. In some embodiments, the electronic device 100 executes a temperature processing strategy using the temperature detected by the temperature sensor 180J.

[0129] The touch sensor 180K, also known as the "touch panel". The touch sensor 180K can be disposed on the display screen 194, and the touch sensor 180K and the display screen 194 form a touch screen, also known as the "touch screen". The touch sensor 180K is used to detect touch operations acting thereon or nearby. The touch sensor can transmit the detected touch operation to the application processor to determine the type of touch event. Visual output related to the touch operation can be provided through the display screen 194. In some other embodiments, the touch sensor 180K can also be disposed on the surface of the electronic device 100, at a different position from that of the display screen 194.

[0130] The bone conduction sensor 180M can acquire vibration signals. In some embodiments, the bone conduction sensor 180M can acquire vibration signals of the vibrating bone mass of the human vocal part. The bone conduction sensor 180M can also contact the human pulse to receive blood pressure pulsation signals. In some embodiments, the bone conduction sensor 180M can also be disposed in the earphone to form a bone conduction earphone. The audio module 170 can analyze the voice signal based on the vibration signal of the vibrating bone mass of the human vocal part acquired by the bone conduction sensor 180M to implement the voice function. The application processor can analyze the heart rate information based on the blood pressure pulsation signal acquired by the bone conduction sensor 180M to implement the heart rate detection function.

[0131] The button 190 includes a power-on button, a volume button, etc. The button 190 can be a mechanical button. It can also be a touch button. The electronic device 100 can receive button inputs to generate key signal inputs related to the user settings and function controls of the electronic device 100.

[0132] The motor 191 can generate vibration prompts. The motor 191 can be used for incoming call vibration prompts and can also be used for touch vibration feedback. For example, touch operations acting on different applications (such as taking pictures, audio playing, etc.) can correspond to different vibration feedback effects. For touch operations acting on different regions of the display screen 194, the motor 191 can also correspond to different vibration feedback effects. Different application scenarios (such as time reminder, receiving information, alarm clock, game, etc.) can also correspond to different vibration feedback effects. The touch vibration feedback effect can also support customization.

[0133] The indicator 192 can be an indicator light, which can be used to indicate the charging state, power change, and can also be used to indicate messages, missed calls, notifications, etc.

[0134] The SIM card interface 195 is used to connect to the SIM card. The SIM card can be inserted into or removed from the SIM card interface 195 to achieve contact with and separation from the electronic device 100. The electronic device 100 can support one or N SIM card interfaces, where N is a positive integer greater than 1. The SIM card interface 195 can support Nano SIM cards, Micro SIM cards, SIM cards, etc. Multiple cards can be inserted into the same SIM card interface 195 at the same time. The types of the multiple cards can be the same or different. The SIM card interface 195 can also be compatible with different types of SIM cards. The SIM card interface 195 can also be compatible with external memory cards. The electronic device 100 interacts with the network through the SIM card to implement functions such as calls and data communication. In some embodiments, the electronic device 100 uses an eSIM, that is, an embedded SIM card. The eSIM card can be embedded in the electronic device 100 and cannot be separated from the electronic device 100.

[0135] The software system of the electronic device 100 can adopt a layered architecture. Figure 2 It is a schematic diagram of the software structure of the electronic device 100 provided by the embodiments of the present application.

[0136] The layered architecture divides the system into several layers, and each layer has a clear role and division of labor. The layers communicate with each other through software interfaces. In some embodiments, the system is divided into five layers, from top to bottom: the application layer, the application framework layer, the hardware abstraction layer, the driver layer, and the hardware layer.

[0137] The application layer can include a series of application program packages. In the embodiments of the present application, the application program packages can include a camera, a gallery, etc.

[0138] The application framework layer provides application programming interfaces (APIs) and programming frameworks for the application programs in the application layer. The application framework layer includes some predefined functions. In the embodiments of the present application, the application framework layer can include a camera access interface, where the camera access interface can include camera management and camera devices. The camera access interface is used to provide application programming interfaces and programming frameworks for camera applications.

[0139] The hardware abstraction layer is an interface layer located between the application framework layer and the driver layer, providing a virtual hardware platform for the operating system. In the embodiments of the present application, the hardware abstraction layer can include a camera hardware abstraction layer and a camera algorithm library.

[0140] Among them, the camera hardware abstraction layer can provide virtual hardware for camera device 1, camera device 2, or more camera devices. The camera algorithm library can include the running code and data for implementing the shooting method provided by the embodiments of the present application.

[0141] The driver layer is the layer between hardware and software. The driver layer includes drivers for various hardware components. The driver layer may include a camera device driver, a digital signal processor driver, an image processor driver, etc.

[0142] Among them, the camera device driver is used to drive the image sensor of the camera to capture images and drive the image signal processor to preprocess the images. The digital signal processor driver is used to drive the digital signal processor to process the images. The image processor driver is used to drive the graphics processor to process the images.

[0143] Next, in combination with the above software structure, the working processes of the software and hardware in the embodiments of the present application when taking pictures using the electronic device 100 will be exemplarily described.

[0144] In response to the user's operation of opening the camera application, such as clicking on the camera application icon, the camera application calls the camera access interface of the application framework layer to start the camera application, and then sends an instruction to start the camera through the camera device (the camera device and / or other camera devices) in the camera hardware abstraction layer. The camera hardware abstraction layer sends this instruction to the camera device driver in the kernel layer. This camera device driver can start the image sensor of the corresponding camera and collect image optical signals through the image sensor. One camera device in the camera hardware abstraction layer corresponds to one image sensor in the hardware layer.

[0145] Then, the image sensor of the camera can transmit the collected image optical signals to the image signal processor for preprocessing to obtain image electrical signals (i.e., the original images, such as RAW images), and transmit the above original images to the camera hardware abstraction layer through the camera device driver.

[0146] The camera hardware abstraction layer can send the original images to the camera algorithm library. The camera algorithm library stores program codes for implementing the method for removing crosstalk artifacts provided in the embodiments of the present application. Based on the digital signal processor and the image processor, the camera algorithm library executes the above codes, enabling the electronic device 100 to execute some or all of the steps of the method for removing crosstalk artifacts provided in the embodiments of the present application.

[0147] The camera algorithm library can send the processed images (such as RGB images) to the camera hardware abstraction layer. Then, the camera hardware abstraction layer can display them. At the same time, the camera algorithm library can also perform various image processing tasks, such as noise reduction, color correction, contrast adjustment, etc., to improve the image quality or perform specific computer vision tasks.

[0148] Optionally, through the interface provided by the camera hardware abstraction layer, the above-mentioned original image or the image processed by the camera algorithm library can be stored in a specific storage unit. Further, in response to the user's operation of opening the gallery application, such as clicking on the gallery application icon and then clicking on the operation to view the image, the gallery application calls the corresponding interface to access the image data in the storage unit, and then calls the image decoding library, etc. to decode the stored image data into an image that can be displayed on the display screen for display in the application.

[0149] It should be noted that in the process of taking pictures using the electronic device 100 in the embodiments of the present application, the original image and the image processed by the camera algorithm library obtained can be static images obtained based on the image sensor of the camera, or dynamic video frames. The embodiments of the present application do not limit this.

[0150] Combined with the above relevant descriptions of the hardware structure and software structure of the electronic device 100, the following introduces the schematic diagrams of some embodiments of the user interface provided by the embodiments of the present application.

[0151] Exemplarily, please refer to Figure 3A , Figure 3A which is a schematic diagram of a user interface for removing crosstalk artifacts in a photographing scenario provided by an embodiment of the present application. As Figure 3A shown, the user turns on the electronic device, so that the display screen of the electronic device displays the desktop of the electronic device, that is, the user interface 31. The user interface 31 may include icons of at least one application program (for example, weather, calendar, mail, settings, app store, note, gallery, phone, short message, browser, and camera, etc.). Among them, the icons of the application programs, as well as the names and positions of the corresponding application programs, can be adjusted according to the user's preferences. The embodiments of the present application do not limit this.

[0152] In the user interface 31, the user can click on the camera 301 control. In response to the operation of clicking on the camera 301 control (i.e., the first operation), the electronic device can display the user interface 32 (i.e., the first interface). In the user interface 32, there is a viewfinder 302 and a shooting control 303. Among them, the viewfinder 302 shows the current picture 304 that the user is shooting (i.e., the first picture). In addition, the user interface 32 may also include a camera switching control 305, which can be used to switch the camera for collecting images between the front camera and the rear camera. Among them, in the picture 304, grid-like crosstalk artifacts caused by crosstalk can be clearly seen. Further, when the user determines to shoot the picture 304 in the viewfinder 302, the user can click on the shooting control 303 (i.e., the second operation). In response to the user's shooting operation (i.e., the first operation), the electronic device switches to the user interface 33 (i.e., the second interface). In the user interface 33, there may be an image preview frame 306, and the image preview frame 306 shows an image 307 (i.e., the fourth image) with crosstalk artifacts that can be used for previewing and corresponds to the picture 304 shown in the viewfinder 302 after the electronic device receives the second operation. When the user clicks on the image preview frame 306 (i.e., the third operation), the electronic device further switches to the user interface 34 (i.e., the third interface). The user interface 34 includes an album display area 308, and the album display area 308 shows an image 309 (i.e., the third image) after removing the crosstalk artifacts. It can be understood that the image 307 and / or the image 309 in the embodiments of the present application can be either a photo or a video, and the embodiments of the present application do not limit this.

[0153] Exemplarily, please refer to Figure 3B , Figure 3B which is another schematic diagram of the user interface for removing crosstalk artifacts in the camera shooting scenario provided by the embodiments of the present application. As Figure 3B shown, for the schematic diagram of the interface when the user turns on the electronic device and starts the camera for shooting and the related description, please refer to Figure 3AThe descriptions of user interface 31 and user interface 32 therein will not be elaborated here. After the user clicks the shooting control 303 (i.e., the second operation), the electronic device switches to user interface 35 (i.e., the second interface). In user interface 35, there is an image preview frame 310. Optionally, there may also be a preview photo saving control 311. Among them, the image displayed in the image preview frame 310 is corresponding to the image 307 (i.e., the fourth image) with crosstalk artifacts that can be used for previewing, which is based on the picture 304 displayed in the viewfinder 302 after the electronic device receives the second operation. Further, when the user clicks the preview photo saving control 311 (i.e., the third operation), the electronic device returns to user interface 33 and then switches to user interface 34. The relevant operations and descriptions of the user in user interface 33 and user interface 34 can be referred to the above Figure 3A related descriptions, which will not be elaborated here. It should be noted that this user interface 35 can be a short-term automatic preview interface that automatically disappears after a few seconds and automatically returns to user interface 33, or it can be an interface that needs to be cancelled after the user manually clicks the return control 312. The embodiments of the present application do not make any limitations.

[0154] Exemplarily, after the user finishes shooting, the user can view the photos or videos without crosstalk artifacts taken through the gallery application. Please refer to Figure 3C , Figure 3C is a schematic diagram of a user interface provided by an embodiment of the present application for viewing an image without crosstalk artifacts after taking a photo. As Figure 3C shown, after the user finishes taking a photo and stores the image 309 (i.e., the third image) without crosstalk artifacts, the electronic device returns to user interface 31. In user interface 31, the user clicks the gallery 313 application. In response to the operation of clicking the gallery 313 application (i.e., the fourth operation), the electronic device switches to user interface 36 (i.e., the fourth interface). In user interface 36, there is a first display area 314 for displaying thumbnails of the images stored in the electronic device. Further, when the user clicks the thumbnail 315 of the image 309 in the first display area 314 (i.e., the fifth operation), the electronic device switches to user interface 37 (i.e., the fifth interface). In user interface 37, there is a second display area 316 for displaying the image 309, so that the user can view the image 309 without crosstalk artifacts after taking a photo.

[0155] It should be noted that Figures 3A - 3CThe schematic diagram of the user interface of the electronic device shown is an exemplary display of the embodiments of the present application. The schematic diagram of the interface of the electronic device can also be in other styles. The number and specific functions of the controls shown in the above user interface are merely exemplary descriptions, and the embodiments of the present application do not limit this. It can be understood that the above operation methods (such as the first operation and / or the second operation, etc.), as well as the display method of the user interface of the electronic device in response to the user's operation, may not be limited to the above operation methods and display methods, and may also include other operation methods (such as drawing a specific shape with a knuckle, or pressing one or more of the volume keys, etc.), and display methods, and the present application does not limit this.

[0156] Based on the foregoing Figures 1 - 2 provided hardware and software structures of the electronic device, and Figures 3A - 3C the relevant descriptions of the user interface embodiments provided, the method steps for removing crosstalk artifacts provided by the embodiments of the present application will be introduced next.

[0157] Exemplarily, please refer to Figure 4A , Figure 4A which is a flowchart example of a method for removing crosstalk artifacts provided by the embodiments of the present application. This method can be applied to the hardware structure of the above Figure 1 electronic device and the Figure 2 software structure of the electronic device, as well as the Figures 3A - 3C user interface provided above. The following will be described with the electronic device as the execution subject.

[0158] Step S4A01: Receive a first operation to start the camera application.

[0159] Step S4A02: In response to the first operation, the display screen displays a first interface.

[0160] Specifically, the first interface includes a viewfinder and a shooting control. The viewfinder is used to display the first picture in real time. Among them, the first picture is an RGB picture of the current shooting scene converted from the RAW picture collected in real time by the image sensor. Exemplarily, the first operation can be Figures 3A - 3C the operation of clicking the camera 301 control in the embodiment. Correspondingly, the first interface can be Figures 3A - 3C the user interface 32 provided in the embodiment, so that the user can observe the current real-time shooting picture (that is, the first picture) through the viewfinder 302 in the user interface 32, so as to facilitate the user to select an appropriate timing to press the camera 301 control to take a picture. In a specific implementation, the first operation may also be other operation methods that can start the camera application, such as pulling down the status bar, or pressing one or more of the volume keys, etc.; the first interface may also be a user interface in other scenarios, and the embodiments of the present application do not limit this.

[0161] Step S4A03: Receive a second operation on the shooting control.

[0162] Step S4A04: In response to the second operation, obtain a first image.

[0163] Specifically, the first image is a RAW image corresponding to the first frame captured by the image sensor after receiving the second operation. When the electronic device receives the second operation on the shooting control, it obtains the RAW image corresponding to the first frame captured by the image sensor (such as a CMOS sensor) after receiving the second operation, thereby obtaining the first image (i.e., the RAW image captured by the image sensor after receiving the second operation). Optionally, the first frame can be a static photo, a dynamic photo, or a video. Therefore, the corresponding first image obtained based on the first frame can be the latest RAW image frame captured by the image sensor after receiving the second operation, or multiple RAW image frames captured by the image sensor after receiving the second operation. Exemplarily, since the pixel arrangement of the image sensor is a multi-Bayer array, that is, multiple pixels of the same color are placed on the size of the original pixels, resulting in a smaller spacing between adjacent pixels. When crosstalk occurs between the same-color pixels of the image sensor during shooting, crosstalk artifacts, such as grid-like crosstalk artifacts, will appear in the RGB image (i.e., the first frame) converted from the RAW image captured by the image sensor in real time by the electronic device. Therefore, there will also be crosstalk artifacts in the RAW image (i.e., the first image) corresponding to the first frame captured by the image sensor after receiving the second operation, reducing the authenticity of the image and affecting the user's shooting experience.

[0164] Exemplarily, the second operation can be Figures 3A - 3C In the embodiment, the operation of clicking the shooting control 303 in the user interface 32. In response to this second operation, the electronic device obtains the RAW image corresponding to the frame captured by the image sensor (such as a CMOS sensor) in the camera after receiving this second operation, that is, the first image. In a specific implementation, the second operation may also be other operation methods that can act on the shooting control, such as drawing a specific shape with a knuckle, or pressing one or more volume buttons, etc. The embodiments of the present application do not limit this.

[0165] Step S4A05: Detect whether the first image includes crosstalk artifacts.

[0166] Specifically, based on the RAW image (i.e., the first image) corresponding to the above-mentioned first frame captured by the image sensor after receiving the second operation, detect whether the first image includes crosstalk artifacts. In one possible implementation, for the above 3A- Figure 3CIn the application scenario of the embodiment, the first image may be a RAW image corresponding to the picture collected by the image sensor after the electronic device receives a second operation of clicking the shooting control. Further, the electronic device may perform a Fourier transform on the RAW image (i.e., the first image) corresponding to the first picture collected by the image sensor after receiving the second operation, to obtain a spectral image (i.e., the fifth image) of the first image converted from the spatial domain to the frequency domain. Further, by determining whether there is a frequency with an intensity greater than a preset value in the spectral image (i.e., the fifth image), it is determined whether the first image includes crosstalk artifacts. When it is detected that there is a frequency with an intensity greater than the preset value in the spectral image (i.e., the fifth image), it is determined that the first image includes crosstalk artifacts (such as grid-like crosstalk artifacts). Thus, through a preset neural network model, the crosstalk artifacts in the area with crosstalk artifacts in the first image can be removed, and the area without crosstalk artifacts is not processed, so as to remove the crosstalk artifacts in the image to restore the authenticity of the image while maximizing the preservation of image clarity, thereby improving the user's shooting experience.

[0167] Step S4A06: If the first image includes crosstalk artifacts, then based on a preset neural network model, correct the first image to obtain a second image.

[0168] Specifically, the neural network model is used to remove the crosstalk artifacts in the area with crosstalk artifacts in the first image; the second image is a RAW image after removing the crosstalk artifacts. Exemplarily, when it is detected that the RAW image (i.e., the first image) corresponding to the first picture collected by the image sensor after the electronic device receives the second operation includes crosstalk artifacts, such as grid-like crosstalk artifacts, further correct the RAW image through a preset neural network model, so as to specifically remove the crosstalk artifacts in the area with crosstalk artifacts in the first image, and obtain a RAW image after removing the crosstalk artifacts (i.e., the second image). Therefore, through the preset neural network model in the embodiment of the present application, the image clarity and authenticity can be restored to the greatest extent, and at the same time, the efficiency of image processing can be improved, and the user's shooting experience can be enhanced.

[0169] Step S4A07: Based on the second image, convert to obtain a third image; the third image is an RGB image after removing the crosstalk artifacts.

[0170] Specifically, based on the RAW image after removing the crosstalk artifacts (i.e., the second image) obtained in the above step S4A07, convert to obtain an RGB image after removing the crosstalk artifacts (i.e., the third image). Optionally, the third image may be a static photo after removing the crosstalk artifacts, or a dynamic photo or video, and the embodiment of the present application is not limited thereto.

[0171] Exemplarily, please refer to Figure 4B , Figure 4B which is a flowchart example of another method for removing crosstalk artifacts provided by an embodiment of the present application. This method can be applied to the hardware structure of the electronic device described above in Figure 1 and the software structure of the electronic device described above in Figure 2 , as well as the user interface provided above in Figures 3A - 3C , and may include the following steps S4B01 - S4B13.

[0172] Step S4B01: Receive a first operation to start the camera application.

[0173] Step S4B02: In response to the first operation, the display screen displays a first interface.

[0174] Among them, the first interface includes a viewfinder and shooting controls. The viewfinder is used to display the first picture in real time. The first picture is an RGB image of the current shooting scene converted from the RAW image collected by the image sensor in real time.

[0175] Step S4B03: Receive a second operation on the shooting control.

[0176] Step S4B04: In response to the second operation, obtain a first image.

[0177] Among them, the first image is the RAW image corresponding to the first picture collected by the image sensor after receiving the second operation.

[0178] Specifically, for the specific descriptions of steps S4B01 - S4B04, please refer to the relevant descriptions of steps S4A01 - S4A04 above, and details will not be repeated here.

[0179] Step S4B05: After receiving the second operation on the shooting control, obtain a fourth image.

[0180] Specifically, the fourth image is the RGB image corresponding to the first picture after receiving the second operation, that is, the RGB image converted from the RAW image collected by the image sensor after the electronic device receives the second operation on the shooting control.

[0181] Exemplarily, the second operation may be Figures 3A - 3CIn the embodiment, for the operation of clicking the shooting control 303 in the user interface 32, after the electronic device receives the second operation, it can obtain the RGB image converted from the RAW image collected by the image sensor, that is, the fourth image. Optionally, the fourth image may be the RGB image corresponding to the first frame of the latest frame after receiving the second operation, or the RGB images corresponding to multiple frames of the first frame after receiving the second operation. It should be noted that there is no clear sequence between step S4B05 and the above step S4B04. It may be performed simultaneously with the above step S4B04, or before or after step S4B04.

[0182] Step S4B06: The display screen displays the second interface; the second interface includes an image preview frame, and the fourth image is displayed in the image preview frame.

[0183] Specifically, the display screen of the electronic device further displays the second interface including the image preview frame, wherein the RGB image corresponding to the first frame (the RGB image converted from the RAW image collected by the image sensor) obtained by the electronic device after receiving the second operation on the shooting control is displayed in the image preview frame. Since the processing operation of removing crosstalk artifacts in the image may take time, when the user completes shooting, the user can immediately view the just-taken photo or video through the image preview frame to initially view the shooting effect. Then, after the electronic device receives the third operation on the image preview frame, the user can directly view the image after removing crosstalk artifacts, thereby avoiding the situation that the shooting experience is affected because the user cannot view the taken photo in time due to waiting too long for the electronic device to process the crosstalk artifacts in the image, and improving the user's shooting experience.

[0184] In a possible implementation manner, the second interface further includes a preview photo saving control, and the third operation includes an operation on the preview photo saving control; or, the third operation includes a viewing operation on the image preview frame. In the embodiment of the present application, when the display screen of the electronic device displays the second interface including the image preview frame, the second interface may further include a preview photo saving control, so that the user can select whether to save the just-taken photo or video; correspondingly, the third operation on the image preview frame may include an operation on the preview photo saving control; or, the third operation may include a viewing operation on the image preview frame, so that the electronic device can, in response to the third operation, display the RGB image after removing crosstalk artifacts (that is, the third image) through the album display frame of the third interface. Through the embodiment of the present application, the user can further select whether to save the just-taken photo or video through the preview photo saving control, making the user's photo-taking operation more intuitive and convenient, thereby improving the user's photo-taking experience.

[0185] Exemplarily, the second interface may be Figures 3A - 3C the user interface 33 in the embodiment. In a specific implementation, the second interface may also be a user interface in other scenarios, which is not limited in the embodiments of the present application.

[0186] Step S4B07: Perform a Fourier transform on the first image to obtain a fifth image; the fifth image is the spectral image of the first image converted from the spatial domain to the frequency domain;

[0187] Step S4B08: Determine whether there is a specific intensity frequency intensity greater than a preset value in the fifth image;

[0188] Step S4B09: If there is a specific frequency intensity greater than the preset value, it is determined that the first image includes crosstalk artifacts.

[0189] Specifically, for the specific description of steps S4B07 - S4B09, reference can be made to the relevant description of step S4A05 above, which will not be elaborated here.

[0190] Step S4B10: Input the first image into a neural network model, and correct the first image through the neural network model to obtain a second image.

[0191] Specifically, the neural network model is used to remove the crosstalk artifacts in the area with crosstalk artifacts in the first image; the second image is the RAW image after removing the crosstalk artifacts. For the description of step S4B10, reference can be made to the relevant description of step S4A06 above, which will not be elaborated here.

[0192] Step S4B11: Based on the second image, convert to obtain a third image; the third image is the RGB image after removing the crosstalk artifacts.

[0193] Specifically, for the specific description of step S4B11, reference can be made to the relevant description of step S4A07 above, which will not be elaborated here.

[0194] Step S4B12: Receive a third operation on the image preview frame.

[0195] Step S4B13: In response to the third operation, the display screen displays a third interface.

[0196] Specifically, the third interface includes an album display area, and the third image is displayed in the album display area. Exemplarily, the third operation may be Figure 3A the operation of clicking on the image preview frame 306 in the embodiment. In response to the third operation, the third interface presented by the electronic device may be Figures 3A - 3CThe user interface 34 provided in the embodiment enables the user to observe the RGB image (i.e., the third image) after crosstalk artifacts are removed through the album display area 308 in the user interface 34. In a specific implementation, the third operation may also be other operation methods, which are not limited in the embodiments of the present application.

[0197] Please refer to Figure 4C , Figure 4C which is a schematic flowchart of a process for viewing an image after artifacts are removed provided by the embodiments of the present application. Optionally, when viewing a photo after crosstalk artifacts are removed after the user finishes shooting, that is, in the above Figure 4A steps S4A01 - S4A07, or Figure 4B after the steps S4B01 - S4B13 described above, the following Figure 4C steps S4C01 - S4C05 may also be included.

[0198] Step S4C01: Store the third image.

[0199] Step S4C02: Receive a fourth operation to start the gallery application.

[0200] Step S4C03: In response to the fourth operation, the display screen displays a fourth interface.

[0201] Among them, the fourth interface includes a first display area, and the first display area is used to display thumbnails of the stored images.

[0202] Step S4C04: Receive a fifth operation on the thumbnail of the third image.

[0203] Step S4C05: In response to the fifth operation, the display screen displays a fifth interface; the fifth interface includes a second display area, and the second display area is used to display the third image.

[0204] Specifically, the electronic device can store the RGB image (i.e., the third image) after crosstalk artifacts are removed. When receiving the fourth operation to start the gallery application, in response to this fourth operation, the display screen of the electronic device displays a fourth interface including a first display area for displaying thumbnails of the stored images; further, when the electronic device receives a fifth operation on the thumbnail of the third image, in response to the fifth operation, the display screen of the electronic device displays a fifth interface including a second display area for displaying the third image. Through the embodiments of the present application, the user can view the photo or video after crosstalk artifacts are removed through the gallery application, enabling the user to more conveniently and quickly view the image after crosstalk artifacts are removed and enhancing the user's shooting experience.

[0205] Exemplarily, the fourth operation may be Figure 3CIn the embodiment, when an operation of clicking on the gallery application is performed on the user interface 31 in response to the fourth operation, the fourth interface presented by the electronic device may be Figures 3A - 3C In the embodiment, the provided user interface 36. Further, the fifth operation may be an operation of clicking on the thumbnail 315 in the first display area 314. At this time, the fifth interface may be that the display screen shows the user interface 37, and the second display area 316 therein may display the image 309. In a specific implementation, the above operation methods (such as the third operation and / or the fourth operation) may also be other operation methods, and the fourth interface and / or the fifth interface may also be other interfaces in other application scenarios. The embodiments of the present application do not limit this

[0206] In a possible implementation manner, the above neural network model is trained based on a sample data set; wherein, the above sample data set includes N pairs of data pairs, N is a positive integer; the above data pairs include a sample image and a target image, the above sample image is a RAW image with crosstalk artifacts, and the above target image is a RAW image obtained by removing the above crosstalk artifacts from the area with the above crosstalk artifacts in the above sample image based on the domain mean compensation method

[0207] In a possible implementation manner, the step of removing crosstalk artifacts from the area with crosstalk artifacts in the sample image based on the domain mean compensation method may include: for the area with the above crosstalk artifacts in the sample image, calculating the mean and variance of the channel values of each pixel in the pixel matrix of the same feature channel; by reducing the size of the variance and determining whether the intensity of the specific frequency is lower than a preset value to determine the adjusted target channel value; based on the above target channel value, obtaining a target image with crosstalk artifacts removed

[0208] Further, for the area with the above crosstalk artifacts in the above sample image, the above crosstalk artifacts are removed based on the domain mean compensation method. Please refer to Figure 5 , Figure 5 is a schematic diagram of removing crosstalk artifacts by a domain mean compensation method provided by the embodiments of the present application

[0209] As Figure 5 shown Figure 5 In (A), it represents the area with the above crosstalk artifacts in the sample image, which includes pixel matrices of multiple different feature channels (such as R channel, G channel, B channel), and the channel values of each pixel are different. Since the pixel matrix of the sample image is obtained based on the multi-Bayer array of the image sensor, for the pixel matrix of the same feature channel in the area with the above crosstalk artifacts in the above sample image, it may be as Figure 5The 2×2 matrix shown in (B) can also be a pixel matrix in other arrangements (such as a 3×3 matrix), and the embodiments of the present application do not limit this. Further, calculate the mean and variance between the pixels of the pixel matrix of the same feature channel. For example Figure 5 the mean of the channel values (such as R1, R2, R3, R4) of each pixel in the R matrix shown in (B) in, and obtain that the channel value of each pixel is R0; calculate the variance between the channel value of each pixel and the average channel value R0, and obtain a variance matrix with channel values R5, R6, R7, R8 respectively. At this time, the original R matrix with channel values R1, R2, R3, R4 respectively can be regarded as the sum of a mean matrix with all channels being R0 and a variance matrix with channel values R5, R6, R7, R8 respectively. Further, by reducing the variance, for example, multiplying the variance of each calculated pixel by a coefficient between 0 and 1 to reduce the variance of each pixel, and calculate the channel values of each pixel of the pixel matrix with a smaller brightness difference based on the reduced variance. Exemplarily, for the pixel matrix in the area with larger crosstalk artifacts, the coefficient multiplied by the pixel variance is smaller, that is, the amplitude of variance reduction is larger; while for the pixel matrix in the area with smaller crosstalk artifacts, the coefficient multiplied by the pixel variance is smaller, that is, the amplitude of variance reduction is smaller. Since the sensor may be more sensitive to the light intensity of a specific frequency, which may cause crosstalk artifacts in the image, this method reduces the variance so that the specific frequency of the pixel matrix after domain mean compensation is lower than the preset value, thereby removing the crosstalk artifacts. Perform domain mean compensation on all pixel matrices in the area with crosstalk artifacts in the sample image through the above method, and determine whether the intensity of the specific frequency of the obtained pixel matrix is lower than the preset value (usually 50HZ - 60HZ). Among them, the specific frequency is usually related to the refresh rate or read frequency of the image sensor. If the specific frequency of the adjusted pixel matrix is lower than the preset value, then obtain the target image without crosstalk artifacts based on the adjusted target channel values. This target image only removes crosstalk artifacts from the area with crosstalk artifacts. Therefore, using this target image to train the neural network can obtain the target image without crosstalk artifacts while minimizing the loss of image clarity to the greatest extent. In this way, a sample data set is made to train the neural network model to obtain the optimal neural network model for removing crosstalk artifacts, so as to improve the image authenticity and thus enhance the user's shooting experience.

[0210] It should be noted that the above Figure 5 is only a possible implementation process of the domain mean compensation algorithm provided by the embodiments of the present application, and its specific implementation process may change based on different types of pixel matrices. The embodiments of the present application do not limit this.

[0211] Exemplarily, please refer to Figure 6 , Figure 6Schematic diagram of training a neural network model based on the U-Net architecture provided by an embodiment of this application. Exemplarily, a sample image with crosstalk artifacts is input as the input picture for inputting into this neural network model, and the target image after removing crosstalk artifacts from the area with crosstalk artifacts through the above domain mean compensation algorithm is used as the output picture of the neural network model. As Figure 6 The U-net network shown has a total of five layers, and performs 4 times of downsampling and 4 times of upsampling on the picture respectively. For example, when the input picture is a single-channel picture of 572×572, it then becomes a 64-channel feature map of 568×568 through 2 consecutive convolutions.

[0212] (1) The first red arrow in the upper left corner represents an operation that reduces the length and width of the feature map to half of the original. The size transformation here is from (568, 568, 64) to (284, 284, 64), and the subsequent two convolutional layers increase the number of channels of the feature map to 128.

[0213] (2) The analysis method of other operations in the left half is the same as that in (1).

[0214] (3) The number of channels of the feature map in the middle is 1024, and then the length and width of the feature map are increased through upsampling (or transposed convolution).

[0215] (4) The gray arrow indicates "copying" the feature map on the left to the feature map on the right, and the way is by channels. For example, the feature map with a size of (280, 280, 128) in the left half of the figure is connected to the feature map on the right (200, 200, 128) through the second gray arrow, and the resulting size is (200, 200, 256).

[0216] (5) Repeating the analysis in (4), it can be concluded that the size of the finally output image is (388, 388, 64).

[0217] (6) Finally, through a 1×1 convolution, the number of channels of the output picture is reduced to 2.

[0218] Optionally, during the training process through the above neural network model, in addition to using the sample image with crosstalk artifacts and the target image after removing crosstalk artifacts as the input picture and output picture of the neural network model, parameters such as variance can also be input to improve the speed and accuracy of model training to achieve the best effect of model training.

[0219] Further based on the neural network model based on the U-Net architecture as Figure 6 shown, the training steps may include but are not limited to the following steps 1-step 4.

[0220] Step 1: Obtain a pair of data pairs from the dataset used to train the above neural network model;

[0221] Step 2: Use the sample image in the data pair as the input to the above neural network model to obtain the corresponding sample corrected image;

[0222] Step 3: Calculate the loss value between the target image and the corresponding sample corrected image according to the loss function;

[0223] Step 4: Adjust each parameter of the above neural network model based on the calculated loss value, and repeat Steps 1 - 4 until the above loss value reaches a preset index, then obtain the trained neural network model.

[0224] Furthermore, each pair of data pairs includes a sample image with crosstalk artifacts and a target image obtained by removing the crosstalk artifacts from the area with crosstalk artifacts in the sample image based on the domain mean compensation method. By training multiple times to adjust each parameter of the neural network model, the neural network model can only remove the crosstalk artifacts in the area with crosstalk artifacts in the sample image and does not process the area without crosstalk artifacts. That is, the corrected image obtained by the neural network model based on the sample image is as close as possible to the target image corresponding to the sample image, so as to obtain a trained neural network model that can be used to remove the crosstalk artifacts in the area with crosstalk artifacts in the first image, so as to remove the crosstalk artifacts through the neural network model to improve the image authenticity and thus enhance the user's shooting experience.

[0225] It should be noted that the above-trained neural network model can be a neural network or a neural network based on other types of architectures; the loss function used to calculate the loss value between the target image and the corresponding sample corrected image can be an absolute value loss function (also known as the L1 loss function) or other loss functions, and the embodiments of the present application are not limited thereto. The above Steps 1 - 4 are only a possible implementation manner of training the above neural network model and do not constitute a specific limitation on the method for removing crosstalk artifacts provided by the embodiments of the present application.

[0226] The above details the method of the embodiments of the present application. It can be understood that in order for each device to implement the corresponding functions, it includes the corresponding hardware structure and / or software module for executing each function. Combining the units and steps of each example described in the embodiments disclosed in this article, the present application can be implemented in the form of hardware or a combination of hardware and computer software. Whether a certain function is executed in the way of hardware or computer software driving hardware depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present application. The devices provided by the embodiments of the present application are introduced below.

[0227] The method of the embodiments of the present application is elaborated in detail above. Next, the devices provided by the embodiments of the present application will be introduced. Exemplarily, please refer to Figure 7 , Figure 7 FIG. 6 is a schematic structural diagram of a device for removing crosstalk artifacts provided by an embodiment of the present application. The device 700 can be applied to an electronic device, and the electronic device includes a display screen and a camera. The camera includes an image sensor. Among them, the pixels of the image sensor are arranged in a multi-Bayer array; the device 700 may include a first receiving unit 701, a first display unit 702, a second receiving unit 703, a first obtaining unit 704, a first detecting unit 705, a first correcting unit 706, and a first converting unit 707. The detailed descriptions of each unit are as follows:

[0228] The first receiving unit 701 is configured to receive a first operation for starting a camera application;

[0229] The first display unit 702 is configured to, in response to the first operation, display a first interface on the display screen; wherein, the first interface includes a viewfinder and a shooting control. The viewfinder is used to display a first picture in real time, and the first picture is an RGB image of the current shooting scene converted from a RAW image collected in real time by the image sensor;

[0230] The second receiving unit 703 is configured to receive a second operation acting on the shooting control;

[0231] The first obtaining unit 704 is configured to, in response to the second operation, obtain a first image; the first image is a RAW image corresponding to the first picture collected by the image sensor after receiving the second operation;

[0232] The first detecting unit 705 is configured to detect whether the first image includes crosstalk artifacts;

[0233] The first correcting unit 706 is configured to, if the first image includes crosstalk artifacts, correct the first image based on a preset neural network model to obtain a second image; wherein, the neural network model is used to remove crosstalk artifacts in the area with crosstalk artifacts in the first image; the second image is a RAW image after removing crosstalk artifacts;

[0234] The first converting unit 707 is configured to convert a third image based on the second image; the third image is an RGB image after removing crosstalk artifacts.

[0235] The apparatus for removing crosstalk artifacts provided by the embodiments of the present application is applied to an electronic device. After the first receiving unit 701 receives the first operation to start the camera application, the first display unit 702 responds to the first operation and displays a first interface including a viewfinder and shooting controls on the display screen of the electronic device. Among them, the viewfinder is used to display in real time the RGB image of the current shooting scene (i.e., the first picture) converted from the RAW image collected in real time by the image sensor. Further, when the second receiving unit 703 receives the second operation acting on the shooting control, the first obtaining unit 704 responds to the second operation and obtains the RAW image corresponding to the first picture collected by the image sensor after receiving the second operation (i.e., the first image). Still further, the first detection unit 705 detects whether the first image includes crosstalk artifacts. If it is detected that the first image includes crosstalk artifacts, the first correction unit 706 corrects the first image based on a preset neural network model to remove the crosstalk artifacts in the area with crosstalk artifacts in the first image, and obtains the RAW image after removing crosstalk artifacts (i.e., the second image). Then, the first conversion unit 707 converts the second image into an RGB image after removing crosstalk artifacts (i.e., the third image). Since the pixel arrangement of the image sensor is a multi-Bayer array, that is, multiple pixels of the same color are placed on the size of the original pixel, the distance between adjacent pixels becomes smaller, and the crosstalk problem between pixels of the same color is aggravated, so that crosstalk artifacts are likely to appear in the RGB image converted by the electronic device based on the RAW image (such as the first image), such as grid-like crosstalk artifacts, thus reducing the authenticity of the captured image. Therefore, through the embodiments of the present application, it is possible to first detect the RAW image corresponding to the first picture collected by the image sensor after the electronic device receives the second operation, that is, to determine whether it includes crosstalk artifacts by detecting the first image. Further, when it is detected that the RAW image (i.e., the first image) obtained by the electronic device based on the image sensor includes crosstalk artifacts, the crosstalk artifacts in only the area with crosstalk artifacts in the first image are removed through a preset neural network model, and the area without crosstalk artifacts in the first image is not processed, and the RAW image after removing crosstalk artifacts (i.e., the second image) is obtained, so that in the RGB image after removing crosstalk artifacts (i.e., the third image) converted based on the RAW image, the area with crosstalk artifacts is improved, and the area without crosstalk artifacts can continue to maintain the previous clarity, so as to greatly restore the authenticity and clarity of the image, thereby enhancing the shooting experience of the user.

[0236] In a possible implementation manner, the apparatus further includes:

[0237] A second acquisition unit, configured to acquire a fourth image after receiving a second operation on the shooting control; the fourth image is an RGB image corresponding to the first picture after receiving the second operation.

[0238] A second display unit, configured to display a second interface on the display screen; the second interface includes an image preview frame, and the fourth image is displayed in the image preview frame.

[0239] A third receiving unit, configured to receive a third operation on the image preview frame.

[0240] A third display unit, configured to display a third interface on the display screen in response to the third operation; the third interface includes an album display area, and the third image is displayed in the album display area.

[0241] In a possible implementation manner, the second interface further includes a preview photo saving control, and the third operation includes an operation on the preview photo saving control; or, the third operation includes a viewing operation on the image preview frame.

[0242] In a possible implementation manner, the device further includes:

[0243] A first storage unit, configured to store the third image.

[0244] A fourth receiving unit, configured to receive a fourth operation to start a gallery application.

[0245] A fourth display unit, configured to display a fourth interface on the display screen in response to the fourth operation; the fourth interface includes a first display area, and the first display area is used to display thumbnails of stored images.

[0246] A fifth receiving unit, configured to receive a fifth operation on the thumbnail of the third image.

[0247] A fifth display unit, configured to display a fifth interface on the display screen in response to the fifth operation; the fifth interface includes a second display area, and the second display area is used to display the third image.

[0248] In a possible implementation manner, the first detection unit 705 is specifically configured to:

[0249] Perform a Fourier transform on the first image to obtain a fifth image; the fifth image is a spectral image of the first image converted from the spatial domain to the frequency domain.

[0250] Determine whether there is a specific frequency intensity greater than a preset value in the fifth image.

[0251] If there exists a specific frequency intensity greater than the preset value, it is determined that the first image includes crosstalk artifacts.

[0252] In a possible implementation manner, the first correction unit 706 is specifically configured to:

[0253] Input the first image into the neural network model, and correct the first image through the neural network model to obtain the second image.

[0254] In a possible implementation manner, the neural network model is trained based on a sample data set; wherein, the sample data set includes N pairs of data pairs, N is a positive integer; each data pair includes a sample image and a target image, the sample image is a RAW image with crosstalk artifacts, and the target image is a RAW image obtained by removing the crosstalk artifacts from the area with the crosstalk artifacts in the sample image based on the domain mean compensation method.

[0255] In a possible implementation manner, removing the crosstalk artifacts from the area with the crosstalk artifacts in the sample image based on the domain mean compensation method includes:

[0256] For the area with the crosstalk artifacts in the sample image, calculate the mean and variance of the channel values of each pixel in the pixel matrix of the same feature channel;

[0257] By reducing the magnitude of the variance, and determining the adjusted target channel value by judging whether the intensity of the specific frequency is lower than the preset value;

[0258] Based on the target channel value, obtain the target image after removing the crosstalk artifacts.

[0259] It should be noted that for the functions of each unit in the device 700 described in the embodiments of the present application, reference may be made to the relevant descriptions in the above method embodiments, and details are not described herein again. It can be understood that the device and method provided in the embodiments of the present application can be implemented in other ways. For example, the system embodiments described above are merely illustrative. For example, the above-mentioned module or unit division is only a logical function division, and there may be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point, the displayed or discussed mutual coupling or direct coupling or communication connection may be through some interfaces, and the indirect coupling or communication connection of the device or unit may be in an electrical, mechanical or other form.

[0260] Exemplarily, please refer to Figure 8 , Figure 8 is a schematic hardware structure diagram of another electronic device provided in the embodiments of the present application. AsFigure 8 As shown, the electronic device 800 includes at least one processor 801 and a memory 802. Among them, the processor 801 is coupled to the memory 802. The coupling in the embodiments of the present application can be a communication connection, which can be electrical or other forms. In addition, the electronic device 800 provided in the embodiments of the present application may further include: a camera 803 and at least one display screen 804( Figure 8 only one is shown), and the camera includes an image sensor 805. The processor 801, the memory 802, the camera 803, and the display screen 804 can be connected through a bus 806. Specifically, the memory 802 is used to store program instructions. The processor 801 is used to call the program instructions stored in the memory 802, so that the electronic device 800 can execute the steps in the control method provided in the embodiments of the present application. The descriptions of its various components and related steps can be referred to the above, and will not be elaborated here.

[0261] It should be noted that the electronic device 800 provided in the embodiments of the present application may include more or fewer components than those shown in the figure, or combine certain components, or split certain components, or different component arrangements. The components shown in the figure can be implemented in hardware, software, or any combination of software and hardware.

[0262] The embodiments of the present application provide a computer-readable storage medium. The computer-readable storage medium stores a computer program, and the computer program is executed by the processor of the above routing device to implement the steps performed by the routing device in the method for restoring a network provided in the embodiments of the present application; or the computer program is executed by the processor of the above electronic device to implement the steps performed by the electronic device in the method for restoring a network provided in the embodiments of the present application.

[0263] The embodiments of the present application provide a computer program. The computer program includes instructions, and the computer program is executed by the processor of the above routing device to implement the steps performed by the routing device in the method for restoring a network provided in the embodiments of the present application; or the computer program is executed by the processor of the above electronic device to implement the steps performed by the electronic device in the method for restoring a network provided in the embodiments of the present application.

[0264] In the above embodiments, the descriptions of each embodiment have their own emphases. For the parts not detailed in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.

[0265] It should be noted that for the foregoing method embodiments, for the sake of simple description, they are all expressed as a series of action combinations. However, those skilled in the art should be aware that this application is not limited by the described action sequence. Because according to this application, certain steps may be performed in other sequences or simultaneously, or certain steps may be omitted, multiple steps may be combined into one step for execution, and / or one step may be decomposed into multiple steps for execution. Secondly, those skilled in the art should also be aware that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily essential to this application. It should also be noted that the features and functions of two or more devices according to the present disclosure may be embodied in one device. Conversely, the features and functions of one device described above may be further divided and embodied by multiple devices.

[0266] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions described in this application are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center in a wired manner (such as coaxial cable, optical fiber, digital subscriber line) or a wireless manner (such as infrared, wireless, microwave, etc.). The computer-readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server or a data center that includes one or more integrated available media. The available medium can be a magnetic medium (for example, a floppy disk, a hard disk, a magnetic tape), an optical medium (for example, a DVD), or a semiconductor medium (for example, a solid-state disk (SSD)), etc.

[0267] Those of ordinary skill in the art can understand all or part of the processes in the above method embodiments. The processes can be completed by relevant hardware instructed by a computer program. The program can be stored in a computer-readable storage medium. When the program is executed, it can include the processes of the above method embodiments. The foregoing storage media include: various media such as ROM or random access memory RAM, magnetic disks, or optical discs that can store program codes.

[0268] In summary, the above description is only an embodiment of the technical solution of this application and is not intended to limit the protection scope of this application. Any modifications, equivalent replacements, improvements, etc. made in accordance with the disclosure of this application shall be included within the protection scope of this application.

Claims

1. A method for removing crosstalk artifacts, characterized in that, it is applied to an electronic device, the electronic device includes a display screen and a camera, the camera includes an image sensor, wherein the pixels of the image sensor are arranged in a multi-Bayer array; the method includes: Receiving a first operation to start the camera application; In response to the first operation, the display screen displays a first interface; wherein, the first interface includes a viewfinder and shooting controls, and the viewfinder is used to display a first picture in real time, and the first picture is an RGB image of the current shooting scene converted from the RAW image collected by the image sensor in real time; Receiving a second operation acting on the shooting control; In response to the second operation, obtaining a first image; the first image is the RAW image corresponding to the first picture collected by the image sensor after receiving the second operation; Detecting whether the first image includes crosstalk artifacts; If the first image includes crosstalk artifacts, then based on a preset neural network model, correcting the first image to obtain a second image; wherein, the neural network model is used to remove the crosstalk artifacts in the area with crosstalk artifacts in the first image; the second image is the RAW image after removing crosstalk artifacts; Based on the second image, converting to obtain a third image; the third image is the RGB image after removing crosstalk artifacts.

2. The method according to claim 1, characterized in that, the method further includes: After receiving the second operation acting on the shooting control, obtaining a fourth image; the fourth image is the RGB image corresponding to the first picture after receiving the second operation; The display screen displays a second interface; the second interface includes an image preview frame, and the fourth image is displayed in the image preview frame; Receiving a third operation on the image preview frame; In response to the third operation, the display screen displays a third interface; the third interface includes an album display area, and the third image is displayed in the album display area.

3. The method according to claim 2, characterized in that, the second interface further includes a preview photo saving control, and the third operation includes an operation acting on the preview photo saving control; or, the third operation includes a viewing operation acting on the image preview frame.

4. The method according to any one of claims 1-3, characterized in that, the method further includes: Storing the third image; Receiving a fourth operation to start the gallery application; In response to the fourth operation, the display screen displays a fourth interface; the fourth interface includes a first display area, and the first display area is used to display thumbnails of stored images; Receiving a fifth operation on the thumbnail of the third image; In response to the fifth operation, the display screen displays a fifth interface; the fifth interface includes a second display area, and the second display area is used to display the third image.

5. The method according to any one of claims 1-4, characterized in that, the detecting whether the first image includes crosstalk artifacts includes: Perform a Fourier transform on the first image to obtain a fifth image; the fifth image is the spectral image of the first image converted from the spatial domain to the frequency domain; Determine whether there is a specific frequency intensity greater than a preset value in the fifth image; If there is a specific frequency intensity greater than the preset value, it is determined that the first image includes crosstalk artifacts.

6. The method according to any one of claims 1-5, wherein, The correcting the first image based on a preset neural network model to obtain a second image includes: Inputting the first image into the neural network model, and correcting the first image through the neural network model to obtain the second image.

7. The method according to any one of claims 1-6, wherein, The neural network model is trained based on a sample data set; wherein, the sample data set includes N pairs of data pairs, N is a positive integer; the data pair includes a sample image and a target image, the sample image is a RAW image with crosstalk artifacts, and the target image is a RAW image obtained by removing the crosstalk artifacts from the area with the crosstalk artifacts in the sample image based on the domain mean compensation method.

8. The method according to claim 7, wherein, The removing the crosstalk artifacts from the area with the crosstalk artifacts in the sample image based on the domain mean compensation method includes: For the area with the crosstalk artifacts in the sample image, calculate the mean and variance of the channel values of each pixel in the pixel matrix of the same feature channel; Reduce the size of the variance, and determine the adjusted target channel value by judging whether the intensity of the specific frequency is lower than a preset value; Based on the target channel value, obtain the target image with the crosstalk artifacts removed.

9. An apparatus for removing crosstalk artifacts, wherein, Applied to an electronic device, the electronic device includes a display screen and a camera, the camera includes an image sensor, wherein the pixels of the image sensor are arranged in a multi-Bayer Bayer array; including: A first receiving unit, configured to receive a first operation for starting a camera application; A first display unit, configured to, in response to the first operation, display a first interface on the display screen; wherein, the first interface includes a viewfinder and a shooting control, and the viewfinder is used to display a first picture, and the first picture is an RGB image of the current shooting scene converted from a RAW image collected in real time by the image sensor; A second receiving unit, configured to receive a second operation acting on the shooting control; A first obtaining unit, configured to, in response to the second operation, obtain a first image; the first image is a RAW image corresponding to the first picture collected by the image sensor after receiving the second operation; A first detecting unit, configured to detect whether the first image includes crosstalk artifacts; A first correction unit, configured to, if crosstalk artifacts are included in the first image, correct the first image based on a preset neural network model to obtain a second image; wherein, the neural network model is used to remove the crosstalk artifacts in the area with crosstalk artifacts in the first image; the second image is a RAW image after removing the crosstalk artifacts. A first conversion unit, configured to convert the second image to obtain a third image; the third image is an RGB image after removing the crosstalk artifacts.

10. An electronic device Characterized in that the electronic device includes a display screen and a camera, the camera includes an image sensor, wherein the pixels of the image sensor are arranged in a multi-Bayer array; the electronic device further includes a memory and a processor, wherein the memory is used to store a computer program, and the processor is used to execute the computer program stored in the memory, so that the electronic device executes the method according to any one of claims 1-8.

11. A computer-readable storage medium Characterized in that the computer-readable storage medium stores a computer program, and the computer program is executed by the processor to implement the method according to any one of claims 1-8.

12. A computer program Characterized in that the computer program includes instructions, and the computer program is executed by a computing device to implement the method according to any one of claims 1-8.

Citation Information

Patent Citations

  • Bayer domain image noise reduction system and method based on non-local mean filtering

    CN110246089A

  • Image display method and device

    CN111741211A

  • Method and apparatus for calibrating image sensor

    CN113301278A

  • Method of calibrating image sensor and device for calibrating image sensor

    US20210266503A1

  • Method and apparatus for processing image artifact by using electronic device

    WO2021141216A1