A method and related apparatus for removing crosstalk artifacts
By using a neural network model to detect and correct RAW images in electronic devices, crosstalk artifacts caused by multi-Bayer arrays are removed, solving the problem of reduced image realism and sharpness and improving the user's shooting experience.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-11-23
- Publication Date
- 2026-03-27
AI Technical Summary
Crosstalk between pixels caused by multi-Bayer arrays in mobile phone cameras leads to crosstalk artifacts when the image signal processor converts the image into an RGB image, affecting image realism and the user's shooting experience.
By using a preset neural network model in electronic devices to detect and correct RAW images acquired by image sensors, crosstalk artifact areas are removed. Only areas with crosstalk artifacts are processed, while the clarity of areas without crosstalk artifacts is preserved, thus achieving the maximum restoration of the realism and clarity of RGB images.
It effectively removes crosstalk artifacts, improves image realism and clarity, enhances the user's shooting experience, and reduces processing time, thus improving the shooting experience.
Smart Images

Figure CN120075625B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of image processing, and in particular, to a method for removing crosstalk artifacts and related equipment. BACKGROUND
[0002] Multi-Bayer array is a technology applied on image sensors, such as Complementary Metal Oxide Semiconductor (CMOS) sensors, which is usually widely used in the camera of electronic devices, such as mobile phone cameras. Due to the size limitation of mobile phone devices, it is difficult to accommodate a larger size sensor, so the multi-Bayer array technology is adopted in mobile phone cameras to introduce multiple color filters of the same color in the single color area of each pixel in the traditional way, thereby increasing the number of perceivable pixels, so as to improve the light sensitivity of the mobile phone camera to improve the shooting performance.
[0003] However, since the multi-Bayer array places multiple pixels of the same color on the size of the original pixel, the distance between adjacent pixels becomes smaller, which exacerbates the crosstalk problem between pixels, resulting in crosstalk artifacts, such as grid-shaped crosstalk artifacts, when the Image Signal Processor (ISP) converts the RAW image into a Red Green Blue (RGB) image.
[0004] The industry currently mainly solves this problem from two levels of hardware and software. At the hardware level, a more advanced CMOS manufacturing process is usually adopted to reduce cross interference; at the software level, a cross interference algorithm is used to calibrate the differences between pixels under gold standard conditions, which can compensate for the differences between pixels to some extent. However, at the hardware level, better processes usually come with higher costs; the software-based solution can only alleviate the problem, but in some extreme cases, such as the presence of a point light source with high contrast with the background (such as the sun or night lights), the same color pixels may still cause crosstalk artifacts in the image due to crosstalk, thereby reducing the authenticity of the image and affecting the user's shooting experience. SUMMARY
[0005] The embodiments of the present application provide a method for removing crosstalk artifacts and related equipment, which can remove crosstalk artifacts in images to improve the user's shooting experience.
[0006] In a first aspect, the embodiments of the present application provide a method for removing crosstalk artifacts, applied to an electronic device, the electronic device comprising a display screen and a camera, the camera comprising an image sensor, wherein the pixels of the image sensor are arranged in a Bayer array; the method can comprise: receiving a first operation for starting a camera application; in response to the first operation, the display screen displays a first interface; wherein the first interface comprises a viewfinder and a shooting control, the viewfinder is used to display a first picture in real time, the first picture is an RGB image of a current shooting scene converted based on a RAW image collected by the image sensor in real time; receiving a second operation acting on the shooting control; in response to the second operation, obtaining a first image; the first image is a RAW image corresponding to the first picture collected by the image sensor after receiving the second operation; detecting whether the first image includes crosstalk artifacts; if the first image includes crosstalk artifacts, correcting the first image based on a preset neural network model to obtain a second image; wherein the neural network model is used to remove crosstalk artifacts in the first image with crosstalk artifact regions; the second image is a RAW image after removing crosstalk artifacts; converting the third image based on the second image; the third image is an RGB image after removing crosstalk artifacts.
[0007] In the embodiment of the present application, when the electronic device detects that the unprocessed RAW image (i.e., the first image) obtained based on the image sensor includes crosstalk artifacts, the first image is corrected by using a preset neural network model for removing crosstalk artifacts in the region of the first image with crosstalk artifacts, to obtain a RAW image after removing crosstalk artifacts (i.e., the second image), and then based on the second RAW image, an RGB image after removing crosstalk artifacts (i.e., the third image) is obtained. Specifically, the display interface (i.e., the first interface) of the display screen of the electronic device after receiving a first operation of starting the camera application includes a viewfinder frame and a shooting control. The viewfinder frame displays a RAW image obtained by real-time collection based on the image sensor (such as a CMOS sensor) in the camera lens, to obtain an RGB image (i.e., the first picture) of the current shooting scene. Further, the electronic device receives a second operation acting on the shooting control, and based on the first picture, obtains the first image (i.e., the RAW image collected by the image sensor corresponding to the first picture after receiving the second operation). Since the pixel arrangement of the image sensor is a multi-Bayer array, i.e., a plurality of same color pixels are arranged on the original pixel, the spacing between adjacent pixels is small, and the crosstalk problem between the same color pixels is aggravated, so that the RGB image converted based on the RAW image (such as the first image) is prone to crosstalk artifacts, such as grid-shaped crosstalk artifacts, thereby reducing the authenticity of the captured image. Therefore, the embodiment of the present application can first detect the RAW image collected by the image sensor corresponding to the first picture after the electronic device receives the second operation, i.e., by detecting the first image, to determine whether it includes crosstalk artifacts. Further, when it is detected that the RAW image (i.e., the first image) obtained based on the image sensor includes crosstalk artifacts, the preset neural network model is used to remove only the crosstalk artifacts in the region of the first image with crosstalk artifacts, without processing the region of the first image without crosstalk artifacts, to obtain a RAW image after removing crosstalk artifacts (i.e., the second image), so that the region with crosstalk artifacts in the RGB image (i.e., the third image) converted based on the RAW image after removing crosstalk artifacts is improved, and the region without crosstalk artifacts can continue to maintain the previous clarity, greatly restoring the authenticity and clarity of the image. Unlike the prior art, which uniformly and roughly removes crosstalk artifacts from the entire image, resulting in a great reduction in the authenticity and clarity of the image, the embodiment of the present application can maximize the protection and restoration of the authenticity and clarity of the image, and improve the user's shooting experience.
[0008] In a possible implementation, the method further includes: after receiving the second operation acting on the shooting control, acquiring a fourth image; the fourth image is an RGB image corresponding to the first image after receiving the second operation; the display screen displays a second interface; the second interface includes an image preview frame, and the fourth image is displayed in the image preview frame; receiving a third operation on the image preview frame; in response to the third operation, the display screen displays a third interface; the third interface includes an album display area, and the third image is displayed in the album display area.
[0009] In the embodiment, after the electronic device receives the second operation acting on the shooting control, the electronic device acquires the RAW image (i.e., the first image) collected by the image sensor after receiving the second operation based on the first image displayed in the viewfinder in the first interface, and also acquires the RGB image (i.e., the fourth image) corresponding to the first image after receiving the second operation. Optionally, because the first image displayed in the viewfinder after receiving the second operation can be a static photo, or a dynamic photo or video, the fourth image can be the RGB image corresponding to the latest first image after receiving the second operation, or the RGB image corresponding to multiple first images after receiving the second operation; correspondingly, the first image acquired based on the first image can be the latest RAW image collected by the image sensor after receiving the second operation, or multiple RAW images collected by the image sensor after receiving the second operation. Further, the display screen of the electronic device displays the second interface including the image preview frame, and the fourth image (i.e., the RGB image corresponding to the latest first image after receiving the second operation, or the RGB image corresponding to multiple first images after receiving the second operation) is displayed in the image preview frame. After receiving the third operation on the image preview frame, the display screen displays the third interface including the album display area, and the RGB image (i.e., the third image) after removing the crosstalk artifacts is displayed in the album display area. According to the embodiment, because the processing operation of removing the crosstalk artifacts in the image can take time, when the user can view the photo or video just taken through the image preview frame immediately after completing the shooting to preliminarily view the shooting effect, the user can directly view the image after removing the crosstalk artifacts after the electronic device receives the third operation on the image preview frame, thereby avoiding the situation that the user cannot view the taken photo in time due to the long time of waiting for the electronic device to process the crosstalk artifacts in the image, and improving the user's shooting experience.
[0010] In a possible implementation, the second interface further includes a preview photo saving control, and the third operation includes an operation on the preview photo saving control; or the third operation includes a viewing operation on the image preview frame.
[0011] According to the embodiment, when the display screen of the electronic device displays the second interface including the image preview frame, the second interface can further include a preview photo saving control, so that the user can select whether to save the just photographed photo or video; accordingly, the third operation on the image preview frame can include an operation on the preview photo saving control; or the third operation can include a viewing operation on the image preview frame, so that the electronic device can display the RGB image (i.e., the third image) after removing the crosstalk artifact through the album display frame of the third interface in response to the third operation. According to the embodiment, the user can further select whether to save the just photographed photo or video through the preview photo saving control, so that the user's photographing operation is more intuitive and convenient, thereby improving the user's photographing experience.
[0012] In a possible implementation, the method further includes: storing the third image; receiving a fourth operation of starting a gallery application; in response to the fourth operation, the display screen displays a fourth interface; the fourth interface includes a first display area for displaying thumbnails of stored images; receiving a fifth operation on the thumbnail of the third image; in response to the fifth operation, the display screen displays a fifth interface; the fifth interface includes a second display area for displaying the third image.
[0013] According to the embodiment, the electronic device can store the RGB image (i.e., the third image) after removing the crosstalk artifact, and when receiving a fourth operation of starting a gallery application, the display screen of the electronic device displays a fourth interface including a first display area for displaying thumbnails of stored images; further, when the electronic device receives a fifth operation on the thumbnail of the third image, the display screen of the electronic device displays a fifth interface including a second display area for displaying the third image. According to the embodiment, the user can view the photo or video after removing the crosstalk artifact through the gallery application, so that the user can more conveniently and quickly view the image after removing the crosstalk artifact, thereby improving the user's photographing experience.
[0014] In a possible implementation, the detecting whether the first image includes the crosstalk artifact can include: performing Fourier transform on the first image to obtain a fifth image; the fifth image is a spectrum image of the first image converted from a spatial domain to a frequency domain; determining whether there is a specific frequency intensity greater than a preset value in the fifth image; and if there is a specific frequency intensity greater than the preset value, determining that the first image includes the crosstalk artifact.
[0015] In the embodiment, the RAW image corresponding to the first picture collected by the image sensor after the electronic device receives the second operation (that is, the first image) is subjected to Fourier transform to obtain a spectrum image (that is, the fifth image) of the first image converted from a spatial domain to a frequency domain. It is determined whether there is a specific frequency intensity greater than a preset value in the spectrum image (that is, the fifth image) to determine whether the first image includes the crosstalk artifact. When it is detected that there is a specific frequency intensity greater than a preset value in the spectrum image (that is, the fifth image), it is determined that the first image includes the crosstalk artifact. Therefore, the preset neural network model can be used to eliminate the crosstalk artifact in the region with the crosstalk artifact in the first image, and the region without the crosstalk artifact is not processed. Therefore, the image clarity is maximally retained, the crosstalk artifact in the image is removed to restore the authenticity of the image, and the user's shooting experience is improved.
[0016] In a possible implementation, the correcting the first image based on the preset neural network model to obtain a second image can include: inputting the first image into the neural network model, correcting the first image by using the neural network model to obtain the second image.
[0017] In the embodiment, the RAW image (that is, the first image) collected by the image sensor after the electronic device receives the second operation is input into the preset neural network model for removing the crosstalk artifact in the region with the crosstalk artifact in the first image. The electronic device can correct the first image by using the neural network model to remove the crosstalk artifact in the region with the crosstalk artifact in the first image, and obtain the RGB image (that is, the second image) after the crosstalk artifact is removed. In the embodiment, the neural network model is used to remove the crosstalk artifact in the region with the crosstalk artifact in the image, so that the image clarity and authenticity are maximally restored, the efficiency of image processing is improved, and the user's shooting experience is improved.
[0018] In a possible implementation, the neural network model is trained based on a sample data set; the sample data set includes N pairs of data, N being a positive integer; each pair of data includes a sample image and a target image, the sample image being a RAW image with crosstalk artifacts, and the target image being a RAW image obtained by removing the crosstalk artifacts from a region with the crosstalk artifacts in the sample image based on a domain mean compensation method.
[0019] The neural network model for removing the crosstalk artifacts in the region with the crosstalk artifacts in the RAW image (that is, the first image) collected by the image sensor after the electronic device receives the second operation in the embodiment of the present application is trained based on a sample data set. The sample data set used to train the neural network model can include N pairs of data, N being a positive integer. Further, each pair of data includes a sample image with crosstalk artifacts and a target image obtained by removing the crosstalk artifacts from a region with the crosstalk artifacts in the sample image based on a domain mean compensation method. By training multiple times to adjust the parameters of the neural network model, the neural network model can only remove the crosstalk artifacts in the region with the crosstalk artifacts in the sample image and does not process the region without the crosstalk artifacts, that is, the corrected image obtained by the neural network model based on the sample image is as close as possible to the target image corresponding to the sample image, so as to obtain the trained neural network model that can be used to remove the crosstalk artifacts in the region with the crosstalk artifacts in the first image, so as to remove the crosstalk artifacts by the neural network model to improve the image authenticity and thus improve the user's shooting experience.
[0020] In a possible implementation, the removing the crosstalk artifacts from the region with the crosstalk artifacts in the sample image based on the domain mean compensation method can include: calculating the mean and variance of the channel values of each pixel in the pixel matrix of the same feature channel for the region with the crosstalk artifacts in the sample image; determining the adjusted target channel value by reducing the size of the variance and judging whether the intensity of the specific frequency is lower than a preset value; and obtaining the target image with the crosstalk artifacts removed based on the target channel value.
[0021] The embodiment of the present application removes the crosstalk artifacts in the sample image with crosstalk artifacts through a domain mean compensation method to obtain a corresponding target image after removing the crosstalk artifacts. Specifically, the mean and variance of the feature channel values of each pixel in the pixel matrix of the same feature channel in the region with crosstalk artifacts in the sample image are calculated, then the size of the variance is reduced, and whether the intensity of the specific frequency is lower than a preset value is judged to adjust the channel values of each pixel in the pixel matrix to obtain the adjusted feature channel values (i.e. target channel values), and further based on the target channel values, the target image corresponding to the sample image after removing the crosstalk artifacts is obtained. Through the embodiment of the present application, the crosstalk artifacts can be removed for the region with crosstalk artifacts based on the sample image with crosstalk artifacts, so as to obtain the target image after removing the crosstalk artifacts while the image clarity is not lost to the greatest extent, so as to make the sample data set by this way to train the neural network model, so as to obtain the optimal neural network model for removing the crosstalk artifacts, so as to improve the image authenticity and thus improve the user's shooting experience.
[0022] In a second aspect, the present application provides a device for removing crosstalk artifacts, applied to an electronic device, the electronic device comprising a display screen and a camera, the camera comprising an image sensor, wherein the pixels of the image sensor are arranged in a Bayer array; which can include:
[0023] A first receiving unit is configured to receive a first operation for starting a camera application;
[0024] A first display unit is configured to, in response to the first operation, display a first interface on the display screen; wherein the first interface comprises a viewfinder and a shooting control, and the viewfinder is configured to display a first picture in real time, and the first picture is an RGB image of a current shooting scene converted based on a RAW image collected by the image sensor in real time;
[0025] A second receiving unit is configured to receive a second operation acting on the shooting control;
[0026] A first obtaining unit is configured to, in response to the second operation, obtain a first image; the first image is a RAW image corresponding to the first picture collected by the image sensor after receiving the second operation;
[0027] A first detection unit is configured to detect whether the first image includes crosstalk artifacts;
[0028] The first correction unit is configured to, if the first image includes crosstalk artifacts, correct the first image based on a preset neural network model to obtain a second image; the neural network model is used to remove crosstalk artifacts in a crosstalk artifact area of the first image; and the second image is a RAW image after the crosstalk artifacts are removed.
[0029] The first conversion unit is configured to convert the second image to obtain a third image; and the third image is an RGB image after the crosstalk artifacts are removed.
[0030] The device for removing crosstalk artifacts provided by the embodiment of the present application is applied to an electronic device. After a first receiving unit receives a first operation of starting a camera application, a first display unit responds to the first operation to display, through a display screen of the electronic device, a first interface including a viewfinder frame and a shooting control. The viewfinder frame is used to display, in real time, an RGB image of a current shooting scene (i.e., a first picture) converted from a RAW image collected in real time based on the image sensor. Further, after a second receiving unit receives a second operation acting on the shooting control, a first obtaining unit responds to the second operation to obtain a RAW image (i.e., a first image) corresponding to the first picture collected by the image sensor after the second operation is received. Further, a first detecting unit detects whether the first image includes crosstalk artifacts. If it is detected that the first image includes crosstalk artifacts, a first correcting unit corrects the first image based on a preset neural network model to remove crosstalk artifacts in a crosstalk artifact region of the first image, to obtain a RAW image after crosstalk artifacts are removed (i.e., a second image), and then a first converting unit converts an RGB image after crosstalk artifacts are removed (i.e., a third image) based on the second image. Since the pixel arrangement of the image sensor is a multi-Bayer array, i.e., a plurality of pixels of the same color are arranged on the original pixel, the distance between adjacent pixels is small, and the crosstalk problem between pixels of the same color is aggravated, so that crosstalk artifacts, such as grid-shaped crosstalk artifacts, are likely to appear in the RGB image converted from the RAW image (e.g., the first image) by the electronic device, thereby reducing the authenticity of the captured image. Therefore, the embodiment of the present application can first detect the RAW image corresponding to the first picture collected by the image sensor after the second operation is received by the electronic device, i.e., by detecting the first image to determine whether the first image includes crosstalk artifacts. Further, when it is detected that the RAW image (i.e., the first image) obtained by the electronic device based on the image sensor includes crosstalk artifacts, the preset neural network model is used to remove only the crosstalk artifacts in the crosstalk artifact region of the first image, and the region of the first image without crosstalk artifacts is not processed, to obtain the RAW image after crosstalk artifacts are removed (i.e., the second image), so that the region with crosstalk artifacts in the RGB image (i.e., the third image) converted based on the RAW image is improved, and the region without crosstalk artifacts can continue to maintain the previous clarity, thereby greatly restoring the authenticity and clarity of the image, and improving the user's shooting experience.
[0031] In a possible implementation, the device further includes:
[0032] The second acquisition unit is configured to acquire a fourth image after receiving a second operation acting on the shooting control; the fourth image is an RGB image corresponding to the first image after receiving the second operation;
[0033] The second display unit is configured to display a second interface on the display screen; the second interface includes an image preview frame, and the fourth image is displayed in the image preview frame;
[0034] The third receiving unit is configured to receive a third operation on the image preview frame;
[0035] The third display unit is configured to display a third interface on the display screen in response to the third operation; the third interface includes a photo album display area, and the third image is displayed in the photo album display area.
[0036] In a possible implementation, the second interface further includes a preview photo saving control, and the third operation includes an operation acting on the preview photo saving control; or the third operation includes a viewing operation on the image preview frame.
[0037] In a possible implementation, the apparatus further includes:
[0038] The first storage unit is configured to store the third image;
[0039] The fourth receiving unit is configured to receive a fourth operation of starting a photo album application;
[0040] The fourth display unit is configured to display a fourth interface on the display screen in response to the fourth operation; the fourth interface includes a first display area, and the first display area is configured to display thumbnails of stored images;
[0041] The fifth receiving unit is configured to receive a fifth operation on the thumbnail of the third image;
[0042] The fifth display unit is configured to display a fifth interface on the display screen in response to the fifth operation; the fifth interface includes a second display area, and the second display area is configured to display the third image.
[0043] In a possible implementation, the first detection unit is specifically configured to:
[0044] perform Fourier transform on the first image to obtain a fifth image; the fifth image is a frequency spectrum image converted from a spatial domain to a frequency domain;
[0045] determine whether a specific frequency intensity in the fifth image is greater than a preset value;
[0046] If there is a specific frequency intensity greater than the preset value, it is determined that the first image includes a crosstalk artifact.
[0047] In a possible implementation, the first correction unit is specifically configured to:
[0048] inputting the first image into the neural network model, correcting the first image through the neural network model, and obtaining the second image.
[0049] In a possible implementation, the neural network model is obtained based on a sample data set; the sample data set includes N pairs of data, N being a positive integer; the data pair includes a sample image and a target image, the sample image being a RAW image with a crosstalk artifact, and the target image being a RAW image obtained by removing the crosstalk artifact from a region with the crosstalk artifact in the sample image based on a domain mean compensation method.
[0050] In a possible implementation, the removing the crosstalk artifact from the region with the crosstalk artifact in the sample image based on the domain mean compensation method includes:
[0051] calculating a mean and a variance of channel values of each pixel in a pixel matrix of the same feature channel for the region with the crosstalk artifact in the sample image;
[0052] determining an adjusted target channel value by reducing the size of the variance and judging whether the intensity of the specific frequency is lower than a preset value;
[0053] obtaining the target image with the crosstalk artifact removed based on the target channel value.
[0054] In a third aspect, an embodiment of the present application provides an electronic device, which can include a memory and a processor, wherein the memory is configured to store a computer program, and the processor is configured to execute the computer program stored in the memory, so that the electronic device executes the method in any one of the second aspect.
[0055] In a fourth aspect, an embodiment of the present application provides a computer readable storage medium, which stores a computer program, and the computer program is executed by the processor to implement the method in any one of the first aspect or the second aspect.
[0056] In a fifth aspect, an embodiment of the present application provides a computer program, which includes instructions, and the computer program is executed by a computing device to implement the method in any one of the first aspect or the second aspect. BRIEF DESCRIPTION OF DRAWINGS
[0057] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the background art, the drawings needed to be used in the embodiments of the present application or the background art will be described below.
[0058] Figure 1 is a schematic diagram of a hardware structure of an electronic device in the prior art.
[0059] Figure 2 is a schematic diagram of a software structure of an electronic device 100 provided by the embodiments of the present application.
[0060] Figure 3A is a schematic diagram of a user interface for removing crosstalk artifacts in a photographing scene provided by the embodiments of the present application.
[0061] Figure 3B is a schematic diagram of another user interface for removing crosstalk artifacts in a photographing scene provided by the embodiments of the present application.
[0062] Figure 3C is a schematic diagram of a user interface for viewing an image after removing crosstalk artifacts provided by the embodiments of the present application.
[0063] Figure 4A is a flowchart of a method for removing crosstalk artifacts provided by the embodiments of the present application.
[0064] Figure 4B is a flowchart of another method for removing crosstalk artifacts provided by the embodiments of the present application.
[0065] Figure 4C is a flowchart of a method for viewing an image after removing crosstalk artifacts provided by the embodiments of the present application.
[0066] Figure 5 is a schematic diagram of a domain mean compensation method for removing crosstalk artifacts provided by the embodiments of the present application.
[0067] Figure 6 is a schematic diagram of a neural network model training based on a U-Net architecture provided by the embodiments of the present application.
[0068] Figure 7 is a schematic diagram of a device for removing crosstalk artifacts provided by the embodiments of the present application.
[0069] Figure 8 is a schematic diagram of a hardware structure of another electronic device provided by the embodiments of the present application. DETAILED DESCRIPTION
[0070] The embodiments of the present application will be described below with reference to the drawings in the embodiments of the present application.
[0071] The terminology used in the description herein is for the purpose of describing particular embodiments only and is not intended to be limiting of the embodiments of the application. As used in the description of the embodiments of the application and the appended claims, the singular forms "a", "an" and "the" are intended to include both the singular and the plural forms, unless the context clearly indicates otherwise. It will also be understood that the term "and / or" as used herein refers to and encompasses any and all possible combinations of one or more of the associated listed items.
[0072] The terms "first", "second", "third", and "fourth", and the like in the description and in the claims of the present application are used for distinguishing between similar objects and not necessarily for describing a particular sequential or chronological order. Moreover, the terms "comprises", "comprising", "includes", "including", "has", "having" and the like, are intended to be inclusive and amount to saying that a process, method, system, product or apparatus that comprises or has a certain characteristic, step, or element includes enough of that characteristic, step, or element to be fairily termed as having that characteristic, step, or element. The terms "comprises", "comprising", "includes", "including", "has", "having" and the like, are also used in the sense of "consisting essentially of" and "consisting of", and the like.
[0073] Reference herein to "an embodiment" means that a particular feature, structure, or characteristic described in connection with the embodiment can be included in at least one embodiment of the application. The appearances of the phrase that invarious places in the specification are not necessarily all referring to the same embodiment, nor are they necessarily mutually exclusive of one another. As will be apparent to those of ordinary skill in the art, embodiments described herein can be combined with other embodiments.
[0074] First, some terms used in the present application are explained to facilitate the understanding of the embodiments of the present application by those skilled in the art.
[0075] (1) Cross-talk: In the field of photography, cross-talk generally refers to unwanted signal interference on an image sensor due to the interaction between adjacent light-sensitive elements.
[0076] (2) Cross-talk artifact: Cross-talk artifact is a disturbance or flaw that appears in an image, often appearing as a grid-like texture or artifact in the image. In digital photography, one of the common cross-talk situations is the interaction between the arrangement structure of the camera's light-sensitive elements (such as the sensor) and the small patterns or textures in the subject. This interaction can cause frequency aliasing, resulting in obvious grid-like artifacts.
[0077] (3) RAW image: refers to an original image that has not been processed in any way, the RAW image retains the original data of each pixel obtained from the sensor, including brightness, color and other related information.
[0078] (4) Channel Value: Refers to the numerical value of a color channel in a digital image. In color images, three primary color channels, red (R), green (G), and blue (B), are commonly used to represent the color of each pixel. Each channel has a numerical value representing the intensity or brightness of the color in that channel.
[0079] (5) Up-sampling: Refers to increasing the sampling rate of a signal or image, i.e., increasing the number of samples. In image processing, up-sampling is often used to enlarge low-resolution images or feature maps to the original resolution for more detailed analysis or comparison with high-resolution images. Common up-sampling methods include nearest neighbor interpolation, bilinear interpolation, and transposed convolution (also known as deconvolution).
[0080] (6) Down-sampling: Refers to reducing the sampling rate of a signal or image, i.e., reducing the number of samples. In image processing, down-sampling is often used to reduce the size of an image, thereby reducing computational cost, storage requirements, or simplifying subsequent processing. Common down-sampling methods include average pooling (AP) and max pooling (MP), where pooling operations extract features of an image by taking the average or maximum value in the image region and reduce the dimensionality of the image.
[0081] (7) Convolutional Neural Network (CNN): A type of deep learning neural network that can effectively capture and learn hierarchical feature representations, making it perform well in image processing and computer vision fields. Due to the translational invariance and parameter sharing characteristics of CNNs, they can effectively process data with spatial hierarchies, mainly used for tasks involving grid-structured data such as image and video recognition, computer vision tasks, etc.
[0082] (8) Convolutional Layer (CL): The convolutional layer is the core part of the CNN. It uses a convolution kernel (filter) to perform convolution operations on the input data, thereby extracting local features in the input data. Convolution operations move the convolution kernel over the input data, calculating the convolution result at each position to generate the output feature map.
[0083] (9) Pooling Layer (PL): The pooling layer is used to reduce the spatial dimension of the feature map, reduce computational complexity, and improve the robustness of the model. Max pooling and average pooling are common pooling operations used to retain the most significant features.
[0084] (10) Loss Function (LF): Defines the difference between the model's output for a given input and the actual label.
[0085] In order to facilitate understanding of the embodiments of the present application, first, an exemplary electronic device provided in the embodiments of the present application is introduced.
[0086] Please refer to Figure 1 , Figure 1 is a hardware structure schematic diagram of an electronic device provided by the embodiments of the present application. The electronic device 100 is a smart terminal device, which can be various types, and the embodiments of the present application do not limit the specific type thereof. For example, the terminal device can be a mobile phone, and can also include a tablet computer, a desktop computer, a desktop computer with a touch-sensitive surface or a touch panel, a laptop, a handheld computer, a notebook computer, a smart screen, a wearable device (such as a smart watch, a smart bracelet, etc.), an augmented reality (AR) device, a virtual reality (VR) device, an artificial intelligence (AI) device, a car machine, a smart earphone, a game machine, and can also be an internet of things (IOT) device or a smart home device such as a smart water heater, a smart lamp, a smart air conditioner, etc.
[0087] Please refer to Figure 1 , and the specific introduction of each component of the electronic device 100 is as follows: Figure 1
[0088] The electronic device 100 can include a processor 110, an external memory interface 120, an internal memory 121, a universal serial bus (USB) interface 130, a charge management module 140, a power management module 141, a battery 142, an antenna 1, an antenna 2, a mobile communication module 150, a wireless communication module 160, an audio module 170, a loudspeaker 170A, a receiver 170B, a microphone 170C, an earphone interface 170D, a sensor module 180, a key 190, a motor 191, an indicator 192, a camera 193, a display screen 194, and a subscriber identification module (SIM) card interface 195, etc. The sensor module 180 can include a pressure sensor 180A, a gyroscope sensor 180B, an air pressure sensor 180C, a magnetic sensor 180D, an acceleration sensor 180E, a distance sensor 180F, a proximity light sensor 180G, a fingerprint sensor 180H, a temperature sensor 180J, a touch sensor 180K, an ambient light sensor 180L, a bone conduction sensor 180M, etc.
[0089] It can be understood that the structural schematic of the embodiments of the present application does not constitute a specific limitation on the electronic device 100. In other embodiments of the present application, the electronic device 100 can include more or fewer components than those shown, or combine certain components, or split certain components, or different arrangement of components. The components shown can be implemented in hardware, software, or a combination of software and hardware.
[0090] The processor 110 can include one or more processing units. For example, the processor 110 can include an application processor (AP), a modem processor, a graphics processing unit (GPU), an image signal processor (ISP), a controller, a memory, a video codec, a digital signal processor (DSP), a baseband processor, and / or a neural-network processing unit (NPU), etc. Different processing units can be independent devices or integrated into one or more processors.
[0091] The controller can be the nerve center and command center of the electronic device 100. The controller can generate operation control signals according to instruction operation codes and timing signals, and complete the control of fetching and executing instructions.
[0092] The memory can also be provided in the processor 110, for storing instructions and data. In some embodiments, the memory in the processor 110 is a cache memory. The memory can save instructions or data that have just been used or are repeatedly used by the processor 110. If the processor 110 needs to use the instructions or data again, it can directly call from the memory. This avoids repeated access and reduces the waiting time of the processor 110, thereby improving the efficiency of the system.
[0093] In some embodiments, the processor 110 can include one or more interfaces. The interfaces can include an inter-integrated circuit (I2C) interface, an inter-integrated circuit sound (I2S) interface, a pulse code modulation (PCM) interface, a universal asynchronous receiver / transmitter (UART) interface, a mobile industry processor interface (MIPI), a general-purpose input / output (GPIO) interface, a subscriber identity module (SIM) interface, and / or a universal serial bus (USB) interface, etc.
[0094] It can be understood that the interface connection relationship between the modules shown in the embodiments of the present application is only illustrative and does not constitute a limitation on the structure of the electronic device 100. In other embodiments of the present application, the electronic device 100 can also use different interface connection methods or a combination of multiple interface connection methods.
[0095] The charging management module 140 is configured to receive charging input from a charger. The charger can be a wireless charger or a wired charger. In some wired charging embodiments, the charging management module 140 can receive charging input from a wired charger through the USB interface 130. In some wireless charging embodiments, the charging management module 140 can receive wireless charging input through a wireless charging coil of the electronic device 100. The charging management module 140 can charge the battery 142 and also supply power to the electronic device through the power management module 141.
[0096] The power management module 141 is configured to connect the battery 142, the charging management module 140, and the processor 110. The power management module 141 receives input from the battery 142 and / or the charging management module 140 to supply power to the processor 110, the internal memory 121, the external memory, the display screen 194, the camera 193, and the wireless communication module 160, etc. The power management module 141 can also be configured to monitor parameters such as battery capacity, battery cycle count, battery health status (leakage, impedance), etc. In other embodiments, the power management module 141 can also be arranged in the processor 110. In other embodiments, the power management module 141 and the charging management module 140 can also be arranged in the same device.
[0097] The wireless communication function of the electronic device 100 can be implemented by the antenna 1, the antenna 2, the mobile communication module 150, the wireless communication module 160, the modem processor, and the baseband processor, etc.
[0098] The antenna 1 and the antenna 2 are used for transmitting and receiving electromagnetic wave signals. Each antenna in the electronic device 100 can be used to cover a single or multiple communication frequency bands. Different antennas can also be multiplexed to improve the utilization of the antennas. For example, the antenna 1 can be multiplexed as a diversity antenna for a wireless local area network. In some other embodiments, the antennas can be used in combination with a tuning switch.
[0099] The mobile communication module 150 can provide a solution including 2G / 3G / 4G / 5G wireless communication applied on the electronic device 100. The mobile communication module 150 can include at least one filter, a switch, a power amplifier, a low noise amplifier (LNA), etc. The mobile communication module 150 can receive electromagnetic waves by the antenna 1, and perform filtering, amplification, etc. on the received electromagnetic waves, and transmit the processed signals to the modem processor for demodulation. The mobile communication module 150 can also amplify the signals modulated by the modem processor, and convert the signals into electromagnetic waves radiated by the antenna 1. In some embodiments, at least part of the functional modules of the mobile communication module 150 can be arranged in the processor 110. In some embodiments, at least part of the functional modules of the mobile communication module 150 can be arranged in the same device as at least part of the modules of the processor 110.
[0100] The modem processor can include a modulator and a demodulator. In some embodiments, the modem processor can be an independent device. In some other embodiments, the modem processor can be independent of the processor 110, and arranged in the same device as the mobile communication module 150 or other functional modules.
[0101] The wireless communication module 160 can provide a solution for wireless communication including wireless local area networks (WLAN) (e.g., wireless fidelity (Wi-Fi) network), Bluetooth (BT), global navigation satellite system (GNSS), frequency modulation (FM), near field communication (NFC), infrared (IR) technology, etc. applied to the electronic device 100. The wireless communication module 160 can be one or more devices that integrate at least one communication processing module. The wireless communication module 160 receives an electromagnetic wave via the antenna 2, frequency-modulates and filters the electromagnetic wave signal, and transmits the processed signal to the processor 110. The wireless communication module 160 can also receive a signal to be transmitted from the processor 110, frequency-modulate it, amplify it, and radiate it as an electromagnetic wave via the antenna 2.
[0102] In some embodiments, the antenna 1 and the mobile communication module 150 of the electronic device 100 are coupled, and the antenna 2 and the wireless communication module 160 are coupled, so that the electronic device 100 can communicate with a network and other devices through wireless communication technology. The wireless communication technology can include global system for mobile communications (GSM), general packet radio service (GPRS), code division multiple access (CDMA), wideband code division multiple access (WCDMA), time-division code division multiple access (TD-SCDMA), long term evolution (LTE), BT, GNSS, WLAN, NFC, FM, and / or IR technology, etc. The GNSS can include a global positioning system (GPS), a global navigation satellite system (GLONASS), a beidu navigation satellite system (BDS), a quasi-zenith satellite system (QZSS), and / or a satellite based augmentation systems (SBAS).
[0103] The electronic device 100 implements a display function through a GPU, a display screen 194, and an application processor, etc. The GPU is a microprocessor for image processing, which is connected to the display screen 194 and the application processor. The GPU is used to perform mathematical and geometric calculations for graphics rendering. The processor 110 can include one or more GPUs, which execute program instructions to generate or change display information.
[0104] The display screen 194 is configured to display images, videos, and the like. The display screen 194 includes a display panel. The display panel can be a liquid crystal display (LCD), an organic light-emitting diode (OLED), an active-matrix organic light-emitting diode (AMOLED), a flex light-emitting diode (FLED), a Miniled, a MicroLed, a Micro-oLed, a quantum dot light emitting diodes (QLED), or the like. In some embodiments, the electronic device 100 can include one or N display screens 194, where N is a positive integer greater than 1.
[0105] The electronic device 100 can implement the photographing function through the ISP, the camera 193, the video codec, the GPU, the display screen 194, and the application processor.
[0106] The ISP is configured to process the data fed back by the camera 193. For example, when taking a photo, the shutter is opened, the light is transmitted to the camera photosensitive element through the lens, the light signal is converted into an electrical signal, and the camera photosensitive element transmits the electrical signal to the ISP for processing to convert it into an image visible to the naked eye. The ISP can also algorithmically optimize the noise and brightness of the image. The ISP can also optimize the exposure, color temperature, and other parameters of the shooting scene. In some embodiments, the ISP can be disposed in the camera 193.
[0107] The camera 193 is configured to capture still images or videos. An object generates an optical image through a lens and projects it onto a photosensitive element. The photosensitive element can be a charge coupled device (CCD) or a complementary metal-oxide-semiconductor (CMOS) phototransistor. The photosensitive element converts the light signal into an electrical signal, which is then transmitted to the ISP to convert it into a digital image signal. The ISP outputs the digital image signal to the DSP for processing. The DSP converts the digital image signal into an image signal in a standard RGB, YUV, or the like format. In some embodiments, the electronic device 100 can include one or N cameras 193, where N is a positive integer greater than 1.
[0108] The digital signal processor is used to process digital signals, in addition to being able to process digital image signals, it can also process other digital signals. For example, when the electronic device 100 selects a frequency point, the digital signal processor is used to perform Fourier transform on the frequency point energy, etc.
[0109] The video codec is used to compress or decompress digital video. The electronic device 100 can support one or more video codecs. In this way, the electronic device 100 can play or record videos in multiple encoding formats, such as: moving picture experts group (MPEG) 1, MPEG 2, MPEG 3, MPEG 4, etc.
[0110] The NPU is a neural-network (NN) calculation processor, which can quickly process input information by drawing on the structure of a biological neural network, such as drawing on the transmission mode between human brain neurons, and can also constantly self-learn. Through the NPU, the electronic device 100 can realize intelligent cognition applications such as image recognition, face recognition, voice recognition, text understanding, etc.
[0111] The external memory interface 120 can be used to connect an external memory card, such as a Micro SD card, to expand the storage capacity of the electronic device 100. The external memory card communicates with the processor 110 through the external memory interface 120 to realize data storage functions. For example, music, video, etc. Files are saved in the external memory card.
[0112] The internal memory 121 can be used to store computer executable program codes, which include instructions. The processor 110 executes various function applications and data processing of the electronic device 100 by running the instructions stored in the internal memory 121. The internal memory 121 can include a program storage area and a data storage area. The program storage area can store an operating system, at least one application program required by a function (such as a sound playing function, an image playing function, etc.), etc. The data storage area can store data created during the use of the electronic device 100 (such as audio data, a phone book, etc.), etc. In addition, the internal memory 121 can include a high-speed random access memory, and can also include a non-volatile memory, such as at least one magnetic disk storage device, a flash memory device, a universal flash storage (UFS), etc.
[0113] The electronic device 100 can realize audio functions through the audio module 170, the speaker 170A, the receiver 170B, the microphone 170C, the earphone interface 170D, and the application processor, etc. For example, music playing, recording, etc.
[0114] The audio module 170 is configured to convert digital audio information into an analog audio signal output, and to convert an analog audio input into a digital audio signal. The audio module 170 can also be configured to encode and decode audio signals. In some embodiments, the audio module 170 can be disposed in the processor 110, or some of the functional modules of the audio module 170 can be disposed in the processor 110.
[0115] The speaker 170A, also referred to as a "loudspeaker", is configured to convert an audio electrical signal into a sound signal. The electronic device 100 can listen to music or listen to a hands-free call through the speaker 170A.
[0116] The receiver 170B, also referred to as a "earpiece", is configured to convert an audio electrical signal into a sound signal. When the electronic device 100 receives a call or a voice message, the user can listen to the voice by holding the receiver 170B close to the ear.
[0117] The microphone 170C, also referred to as a "microphone", "sound collector", is configured to convert a sound signal into an electrical signal. When making a call or sending a voice message, the user can make a sound by holding the mouth close to the microphone 170C, and input the sound signal into the microphone 170C. The electronic device 100 can be provided with at least one microphone 170C. In other embodiments, the electronic device 100 can be provided with two microphones 170C, in addition to collecting sound signals, noise reduction functions can also be realized. In other embodiments, the electronic device 100 can also be provided with three, four or more microphones 170C, to realize the collection of sound signals, noise reduction, and also to identify the source of the sound, to realize the function of directional recording, etc.
[0118] The earphone interface 170D is configured to connect a wired earphone. The earphone interface 170D can be a USB interface 130, or a 3.5mm open mobile terminal platform (OMTP) standard interface, a cellular telecommunications industry association of the USA (CTIA) standard interface.
[0119] The pressure sensor 180A is configured to sense a pressure signal and convert the pressure signal into an electrical signal. In some embodiments, the pressure sensor 180A can be disposed on the display screen 194. The pressure sensor 180A can be of various types, such as a resistive pressure sensor, an inductive pressure sensor, a capacitive pressure sensor, etc. The capacitive pressure sensor can include at least two parallel plates of conductive material. When a force is applied to the pressure sensor 180A, the capacitance between the electrodes changes. The electronic device 100 determines the intensity of the pressure based on the change in capacitance. When a touch operation is applied to the display screen 194, the electronic device 100 detects the intensity of the touch operation based on the pressure sensor 180A. The electronic device 100 can also calculate the position of the touch based on the detection signal of the pressure sensor 180A. In some embodiments, touch operations applied to the same touch position but with different touch operation intensities can correspond to different operation instructions. For example, when a touch operation with a touch operation intensity less than a first pressure threshold is applied to a short message application icon, an instruction to view short messages is executed. When a touch operation with a touch operation intensity greater than or equal to the first pressure threshold is applied to the short message application icon, an instruction to create a new short message is executed.
[0120] The gyroscope sensor 180B can be configured to determine the motion posture of the electronic device 100.
[0121] The barometric pressure sensor 180C is configured to measure barometric pressure.
[0122] The magnetic sensor 180D includes a Hall sensor.
[0123] The acceleration sensor 180E can detect the magnitude of acceleration of the electronic device 100 in various directions (typically three axes). When the electronic device 100 is stationary, the acceleration sensor 180E can detect the magnitude and direction of gravity. The acceleration sensor 180E can also be used to identify the posture of the electronic device 100 and applied to switching between landscape and portrait modes, pedometers, etc.
[0124] The distance sensor 180F is configured to measure distance.
[0125] The proximity light sensor 180G can include, for example, a light emitting diode (LED) and a light detector, such as a photodiode.
[0126] The ambient light sensor 180L is configured to sense ambient light intensity.
[0127] The fingerprint sensor 180H is configured to acquire a fingerprint. The electronic device 100 can use the acquired fingerprint characteristics to implement fingerprint unlocking, access to application locks, fingerprint photographing, fingerprint answering calls, etc.
[0128] The temperature sensor 180J is configured to detect temperature. In some embodiments, the electronic device 100 uses the temperature detected by the temperature sensor 180J to implement a temperature processing strategy.
[0129] Touch sensor 180K, also referred to as "touch panel". Touch sensor 180K can be disposed on display screen 194, and touch screen, also referred to as "touch panel", can be formed by touch sensor 180K and display screen 194. Touch sensor 180K is configured to detect touch operations applied thereto or in the vicinity thereof. Touch sensor 180K can transmit the detected touch operation to the application processor to determine the type of touch event. Visual output related to the touch operation can be provided through display screen 194. In other embodiments, touch sensor 180K can also be disposed on the surface of electronic device 100, which is different from the position of display screen 194.
[0130] Bone conduction sensor 180M can obtain vibration signals. In some embodiments, bone conduction sensor 180M can obtain vibration signals of the bone block of the human body sound part. Bone conduction sensor 180M can also contact the human body pulse to receive blood pressure pulsation signals. In some embodiments, bone conduction sensor 180M can also be disposed in the earphone to form a bone conduction earphone. Audio module 170 can analyze voice signals based on the vibration signals of the bone block of the human body sound part obtained by bone conduction sensor 180M to realize voice functions. The application processor can analyze heart rate information based on the blood pressure pulsation signals obtained by bone conduction sensor 180M to realize heart rate detection functions.
[0131] Keys 190 include power on / off keys, volume keys, and the like. Keys 190 can be mechanical keys. They can also be touch keys. Electronic device 100 can receive key inputs and generate key signal inputs related to user settings and function control of electronic device 100.
[0132] Motor 191 can generate vibration prompts. Motor 191 can be used for incoming call vibration prompts and also for touch vibration feedback. For example, touch operations applied to different applications (such as taking pictures, playing audio, and the like) can correspond to different vibration feedback effects. Touch operations applied to different regions of display screen 194 can also correspond to different vibration feedback effects. Different application scenarios (such as time reminders, received messages, alarms, games, and the like) can also correspond to different vibration feedback effects. Touch vibration feedback effects can also be customizable.
[0133] Indicator 192 can be an indicator light, which can be used to indicate charging status, power changes, and also to indicate messages, missed calls, notifications, and the like.
[0134] The SIM card interface 195 is configured to connect a SIM card. The SIM card can be inserted into or pulled out of the SIM card interface 195 to realize contact and separation with the electronic device 100. The electronic device 100 can support one or N SIM card interfaces, where N is a positive integer greater than 1. The SIM card interface 195 can support a Nano SIM card, a Micro SIM card, a SIM card, and the like. The same SIM card interface 195 can simultaneously insert multiple cards. The types of the multiple cards can be the same or different. The SIM card interface 195 can also be compatible with different types of SIM cards. The SIM card interface 195 can also be compatible with external storage cards. The electronic device 100 interacts with a network through the SIM card to realize functions such as call and data communication. In some embodiments, the electronic device 100 uses an eSIM, that is, an embedded SIM card. The eSIM card can be embedded in the electronic device 100 and cannot be separated from the electronic device 100.
[0135] The software system of the electronic device 100 can adopt a layered architecture, Figure 2 FIG. 1 is a schematic diagram of a software structure of the electronic device 100 provided in an embodiment of the present application.
[0136] The layered architecture divides the system into several layers, each of which has a clear role and division of labor. Layers communicate with each other through software interfaces. In some embodiments, the system is divided into five layers, from top to bottom, an application layer, an application framework layer, a hardware abstraction layer, a driver layer, and a hardware layer.
[0137] The application layer can include a series of application packages. In an embodiment of the present application, the application package can include a camera, a gallery, and the like.
[0138] The application framework layer provides application programming interfaces (APIs) and programming frameworks for the application programs of the application layer. The application framework layer includes some pre-defined functions. In an embodiment of the present application, the application framework layer can include a camera access interface, where the camera access interface can include camera management and camera devices. The camera access interface is configured to provide application programming interfaces and programming frameworks for camera applications.
[0139] The hardware abstraction layer is an interface layer between the application framework layer and the driver layer, and provides a virtual hardware platform for the operating system. In an embodiment of the present application, the hardware abstraction layer can include a camera hardware abstraction layer and a camera algorithm library.
[0140] The camera hardware abstraction layer can provide virtual hardware of a camera device 1, a camera device 2, or more camera devices. The camera algorithm library can include running code and data for implementing the photographing method provided in an embodiment of the present application.
[0141] The driving layer is a layer between hardware and software. The driving layer includes drivers of various hardware. The driving layer can include a camera device driver, a digital signal processor driver, and an image processor driver, etc.
[0142] The camera device driver is used to drive the image sensor of the camera to collect images and drive the image signal processor to pre-process the images. The digital signal processor driver is used to drive the digital signal processor to process the images. The image processor driver is used to drive the image processor to process the images.
[0143] The following illustrates the working process of the software and hardware of the embodiment of the present application when taking a picture by using the electronic device 100 in combination with the above software structure.
[0144] In response to the operation of the user opening the camera application, such as the operation of clicking the camera application icon, the camera application calls the camera access interface of the application framework layer, starts the camera application, and then sends an instruction of starting the camera to a camera device (the camera device and / or other camera device) in the camera hardware abstraction layer by calling the camera device. The camera hardware abstraction layer sends the instruction to the camera device driver in the kernel layer. The camera device driver can start the image sensor of the corresponding camera and collect image light signals through the image sensor. One camera device in the camera hardware abstraction layer corresponds to one image sensor in the hardware layer.
[0145] Then, the image sensor of the camera can transmit the collected image light signals to the image signal processor for pre-processing to obtain image electrical signals (i.e., original images, such as RAW images), and transmit the original images to the camera hardware abstraction layer through the camera device driver.
[0146] The camera hardware abstraction layer can send the original images to the camera algorithm library. The camera algorithm library stores program codes for implementing the method for removing crosstalk artifacts provided in the embodiment of the present application. Based on the digital signal processor and the image processor, the camera algorithm library executes the above codes, so that the electronic device 100 can execute part or all of the steps of the method for removing crosstalk artifacts provided in the embodiment of the present application.
[0147] The camera algorithm library can send the processed images (such as RGB images) to the camera hardware abstraction layer. Then, the camera hardware abstraction layer can display them. Meanwhile, the camera algorithm library can also execute various image processing tasks, such as noise reduction, color correction, contrast adjustment, etc., to improve the image quality or perform specific computer vision tasks.
[0148] Optionally, through the interface provided by the camera hardware abstraction layer, the original image or the image processed by the camera algorithm library can be stored in a specific storage unit. Furthermore, in response to the user's operation of opening the gallery application, such as clicking the gallery application icon and then clicking to view the image, the gallery application calls the corresponding interface to access the image data in the storage unit, and then calls the image decoding library to decode the stored image data into an image that can be displayed on the screen for display in the application.
[0149] It should be noted that, in the process of taking pictures using the electronic device 100 in this application embodiment, the original image obtained and the image processed by the camera algorithm library can be a static image obtained based on the camera's image sensor or a dynamic video frame. This application embodiment does not limit this.
[0150] Based on the above description of the hardware and software structure of the electronic device 100, the following section presents schematic diagrams of some user interface embodiments provided in this application.
[0151] For example, please see Figure 3A , Figure 3A This is a schematic diagram of a user interface for removing crosstalk artifacts in a photography scene, provided as an embodiment of this application. Figure 3A As shown, when a user turns on an electronic device, the device's screen displays the desktop, i.e., user interface 31. User interface 31 may include icons for at least one application (e.g., weather, calendar, email, settings, app store, notes, gallery, phone, text messages, browser, and camera). The application icons, their names, and their locations can be adjusted according to user preferences; this embodiment does not limit this.
[0152] In the user interface 31, the user can click the camera 301 control, and in response to the operation (i.e., a first operation) of clicking the camera 301 control, the electronic device can display a user interface 32 (i.e., a first interface). In the user interface 32, a viewfinder 302 and a shooting control 303 are included, where the viewfinder 302 displays a picture 304 (i.e., a first picture) that the user is currently shooting; in addition, the user interface 32 can also include a conversion camera control 305, which can be used to switch the camera for collecting images between a front camera and a rear camera. In the picture 304, a grid-shaped crosstalk artifact caused by crosstalk can be clearly seen. Further, when the user determines to shoot the picture 304 in the viewfinder 302, the shooting control 303 can be clicked (i.e., a second operation), and in response to the shooting operation (i.e., a first operation) of the user, the electronic device enters a user interface 33 (i.e., a second interface), which can include an image preview frame 306, where the image preview frame 306 displays an image 307 (i.e., a fourth image) with crosstalk artifacts that can be used for preview, which corresponds to the picture 304 displayed in the viewfinder 302 based on the electronic device receiving the second operation. When the user clicks the image preview frame 306 (i.e., a third operation), the electronic device further enters a user interface 34 (i.e., a third interface), which includes a photo album display area 308, where the photo album display area 308 displays an image 309 (i.e., a third image) after removing the crosstalk artifact. It can be understood that the image 307 and / or the image 309 in the embodiments of the present application can be a photo or a video, which is not limited in the embodiments of the present application.
[0153] Exemplarily, please refer to Figure 3B , Figure 3B Another user interface schematic diagram provided by the embodiments of the present application for removing crosstalk artifacts in a shooting scene is provided. As shown in Figure 3B , the user opens the electronic device and starts the camera to shoot the interface schematic diagram and the related description, which can be seen in Figure 3AThe description of the user interface 31 and the user interface 32 in the foregoing is not repeated here. When the user clicks the shooting control 303 (i.e., a second operation), the electronic device enters a user interface 35 (i.e., a second interface), which includes an image preview frame 310 and optionally a preview photo saving control 311. The image preview frame 310 displays the image 307 (i.e., a fourth image) with crosstalk artifacts that can be used for preview, which corresponds to the picture 304 displayed in the viewfinder frame 302 based on the second operation received by the electronic device. Further, when the user clicks the preview photo saving control 311 (i.e., a third operation), the electronic device returns to the user interface 33 and then enters a user interface 34. The related operations and descriptions of the user interface 33 and the user interface 34 can be referred to the foregoing Figure 3A related description is not repeated here. It should be noted that the user interface 35 can be a short automatic preview interface that automatically disappears after a few seconds and automatically returns to the user interface 33, or can be an interface that needs to be manually clicked by the user to return the control 312 before it can be canceled. The present application embodiment is not limited.
[0154] For example, after the user completes the shooting, the user can view the photographed image or video with the crosstalk artifacts removed through the gallery application. Please refer to Figure 3C , Figure 3C A user interface schematic diagram for viewing the image with the crosstalk artifacts removed after the shooting is completed is provided in the present application embodiment. As shown in Figure 3C After the user completes the shooting and stores the image 309 (i.e., a third image) with the crosstalk artifacts removed, the electronic device returns to the user interface 31. In the user interface 31, the user clicks the gallery 313 application. In response to the operation of clicking the gallery 313 application (i.e., a fourth operation), the electronic device enters a user interface 36 (i.e., a fourth interface), which includes a first display area 314 for displaying the thumbnails of the images stored in the electronic device. Further, the user clicks the thumbnail 315 of the image 309 in the first display area 314 (i.e., a fifth operation), and the electronic device enters a user interface 37 (i.e., a fifth interface), which includes a second display area 316 for displaying the image 309, so that the user can view the image 309 with the crosstalk artifacts removed after the shooting is completed.
[0155] It should be noted that Figures 3A-3CThe user interface diagram of the electronic device shown is an exemplary display of the embodiments of the present application, and the interface diagram of the electronic device can also be in other styles. The number and specific functions of the controls displayed by the user interface are merely exemplary and the embodiments of the present application are not limited in this regard. It can be understood that the above-mentioned operation modes (such as the first operation and / or the second operation, etc.) and the display mode of the user interface of the electronic device for user operation can not be limited to the above-mentioned operation modes and display modes, and can also include other operation modes (such as drawing a specific shape by knuckle, or pressing one or more of the volume buttons, etc.) and display modes, which are not limited by the present application.
[0156] Based on the foregoing Figures 1-2 The hardware and software structure of the electronic device provided, and Figures 3A-3C The related description of the user interface embodiments provided, and the steps of the method for removing crosstalk artifacts provided by the embodiments of the present application are introduced.
[0157] Exemplarily, please refer to Figure 4A , Figure 4A A flowchart of a method for removing crosstalk artifacts provided by the embodiments of the present application, which can be applied to the hardware structure of the electronic device described above Figure 1 and the software structure of the electronic device described above Figure 2 , and the user interface provided above Figures 3A-3C , which is described below with the electronic device as the execution subject.
[0158] Step S4A01: receiving a first operation for starting a camera application.
[0159] Step S4A02: in response to the first operation, the display screen displays a first interface.
[0160] Specifically, the first interface includes a viewfinder and a shooting control, and the viewfinder is used to display a first picture in real time, wherein the first picture is a RAW image collected by the image sensor in real time, and an RGB image of the current shooting scene converted therefrom. Exemplarily, the first operation can be Figures 3A-3C the operation of clicking the camera 301 control in the embodiments, and correspondingly, the first interface can be Figures 3A-3C the user interface 32 provided in the embodiments, so that the user can observe the picture (i.e. the first picture) of the current real-time shooting through the viewfinder 302 in the user interface 32, in order to help the user select an appropriate opportunity to press the camera 301 control to take a picture. In a specific implementation, the first operation can also be other operation modes that can start the camera application, such as pulling down the status bar, or pressing one or more of the volume buttons, etc.; and the first interface can also be a user interface in other scenes, which are not limited by the embodiments of the present application.
[0161] Step S4A03: receiving a second operation acting on the shooting control.
[0162] Step S4A04: obtaining a first image in response to the second operation.
[0163] Specifically, the first image is a RAW image corresponding to the first frame captured by the image sensor after receiving the second operation. After receiving the second operation acting on the shooting control, the electronic device obtains a RAW image corresponding to the first frame captured by the image sensor (e.g., a CMOS sensor) after receiving the second operation, and obtains the first image (i.e., the RAW image captured by the image sensor after receiving the second operation). Optionally, the first frame can be a still photo, or a dynamic photo or video. Therefore, the corresponding first image obtained based on the first frame can be the RAW image of the latest frame captured by the image sensor after receiving the second operation, or the RAW image of multiple frames captured by the image sensor after receiving the second operation. Illustratively, since the pixel arrangement of the image sensor is a multi-Bayer array, i.e., multiple same-color pixels are arranged on the original pixel, resulting in a smaller distance between adjacent pixels. When the same-color pixels of the image sensor are crosstalked during shooting, the RGB image (i.e., the first frame) converted by the electronic device based on the RAW image captured by the image sensor in real time will have crosstalk artifacts, such as grid-shaped crosstalk artifacts. Therefore, the RAW image corresponding to the first frame (i.e., the first image) captured by the image sensor after receiving the second operation will also have crosstalk artifacts, reducing the authenticity of the image and affecting the user's shooting experience.
[0164] Illustratively, the second operation can be Figures 3A-3C In the embodiment, the operation of clicking the shooting control 303 in the user interface 32, in response to the second operation, the electronic device obtains a RAW image corresponding to the frame captured by the image sensor (e.g., a CMOS sensor) in the camera after receiving the second operation, i.e., the first image. In a specific implementation, the second operation can also be other operation modes that can act on the shooting control, such as drawing a specific shape with the knuckle, or pressing one or more of the volume buttons, etc. The embodiments of the present application do not limit this.
[0165] Step S4A05: detecting whether the first image includes crosstalk artifacts.
[0166] Specifically, based on the RAW image corresponding to the above-mentioned first frame (i.e., the first image) captured by the image sensor after receiving the second operation, it is detected whether the first image includes crosstalk artifacts. In one possible implementation, for the above-mentioned 3A- Figure 3CIn the application scenario in the embodiment, the first image can be a RAW image corresponding to the picture collected by the image sensor after the electronic device receives the second operation of clicking the shooting control. Further, the electronic device can perform Fourier transform on the RAW image (i.e., the first image) corresponding to the first picture collected by the image sensor after receiving the second operation to obtain a frequency spectrum image (i.e., the fifth image) of the first image converted from the spatial domain to the frequency domain. Further, whether the first image includes crosstalk artifacts is determined by judging whether there is a frequency with an intensity greater than a preset value in the frequency spectrum image (i.e., the fifth image). When it is detected that the intensity of a frequency in the frequency spectrum image (i.e., the fifth image) is greater than the preset value, it is determined that the first image includes crosstalk artifacts (for example, grid-shaped crosstalk artifacts). Therefore, the crosstalk artifacts in the crosstalk artifact area of the first image can be removed by using the preset neural network model, and the areas without crosstalk artifacts are not processed, so that the image clarity is maximally retained, the crosstalk artifacts in the image are removed to restore the authenticity of the image, and the user's shooting experience is improved.
[0167] Step S4A06: If the first image includes crosstalk artifacts, correcting the first image based on the preset neural network model to obtain a second image.
[0168] Specifically, the neural network model is used to remove the crosstalk artifacts in the crosstalk artifact area of the first image. The second image is a RAW image after the crosstalk artifacts are removed. For example, when it is detected that the RAW image (i.e., the first image) corresponding to the first picture collected by the image sensor after the electronic device receives the second operation includes crosstalk artifacts, such as grid-shaped crosstalk artifacts, the RAW image is further corrected by using the preset neural network model, so that the crosstalk artifacts in the crosstalk artifact area of the first image are removed, and a RAW image after the crosstalk artifacts are removed (i.e., the second image) is obtained. Therefore, by using the preset neural network model in the embodiment, the image clarity and authenticity can be maximally restored, the efficiency of image processing can be improved, and the user's shooting experience can be improved.
[0169] Step S4A07: converting a third image based on the second image; the third image is a RGB image after the crosstalk artifacts are removed.
[0170] Specifically, the RAW image after the crosstalk artifacts are removed (i.e., the second image) obtained in the above step S4A07 is converted to obtain a RGB image after the crosstalk artifacts are removed (i.e., the third image). Alternatively, the third image can be a still photo after the crosstalk artifacts are removed, or a dynamic photo or a video, which is not limited in the embodiment.
[0171] Exemplarily, please refer to Figure 4B , Figure 4B Another flowchart of a method for removing crosstalk artifacts provided by the embodiments of the present application is shown in FIG. 4B, which can be applied to the hardware structure of the electronic device of Figure 1 , the software structure of the electronic device of Figure 2 , and the user interface provided by the method of Figures 3A-3C . The method can include the following steps S4B01-S4B13.
[0172] Step S4B01: receiving a first operation for starting a camera application.
[0173] Step S4B02: in response to the first operation, displaying a first interface on the screen.
[0174] The first interface includes a viewfinder and a shooting control. The viewfinder is used to display a first picture in real time. The first picture is an RGB image of a current shooting scene converted from a RAW image collected by an image sensor in real time.
[0175] Step S4B03: receiving a second operation on the shooting control.
[0176] Step S4B04: in response to the second operation, obtaining a first image.
[0177] The first image is a RAW image corresponding to the first picture collected by the image sensor after the second operation is received.
[0178] Specifically, the specific description of steps S4B01-S4B04 can refer to the related description of steps S4A01-S4A04, which will not be repeated here.
[0179] Step S4B05: obtaining a fourth image after the second operation on the shooting control is received.
[0180] Specifically, the fourth image is an RGB image corresponding to the first picture, that is, an RGB image converted from a RAW image collected by the image sensor after the second operation on the shooting control is received.
[0181] Exemplarily, the second operation can be Figures 3A-3CIn the embodiment, when the user clicks the shooting control 303 in the user interface 32, the electronic device can acquire the RGB image converted from the RAW image collected by the image sensor, i.e., the fourth image, after receiving the second operation. Alternatively, the fourth image can be the RGB image corresponding to the latest first frame after receiving the second operation, or the RGB image corresponding to multiple first frames after receiving the second operation. It should be noted that there is no specific sequence between the step S4B05 and the step S4B04 described above, and the step S4B05 can be performed simultaneously with the step S4B04, or before or after the step S4B04.
[0182] In step S4B06, the display screen displays a second interface, and the second interface includes an image preview frame, and the fourth image is displayed in the image preview frame.
[0183] Specifically, the display screen of the electronic device further displays a second interface including an image preview frame, and the fourth image corresponding to the first frame after receiving the second operation acting on the shooting control is displayed in the image preview frame. Since the operation of removing the crosstalk artifacts in the image can take time, the user can immediately view the just-shot photo or video through the image preview frame to preliminarily view the shooting effect after the shooting is completed, and then the user can directly view the image after the crosstalk artifacts are removed when the electronic device receives the third operation acting on the image preview frame, thereby avoiding the situation that the user cannot timely view the shot photo due to the long time for the electronic device to process the crosstalk artifacts in the image, and improving the shooting experience of the user.
[0184] In a possible implementation, the second interface further includes a preview photo saving control, and the third operation includes an operation acting on the preview photo saving control, or the third operation includes a viewing operation acting on the image preview frame. In the embodiment of the present application, when the display screen of the electronic device displays the second interface including the image preview frame, the second interface can further include a preview photo saving control, so that the user can select whether to save the just-shot photo or video; correspondingly, the third operation acting on the image preview frame can include an operation acting on the preview photo saving control, or the third operation can include a viewing operation acting on the image preview frame, so that the electronic device can display the RGB image after the crosstalk artifacts are removed (i.e., the third image) through the album display frame of the third interface in response to the third operation. Through the embodiment of the present application, the user can further select whether to save the just-shot photo or video through the preview photo saving control, so that the user's shooting operation is more intuitive and convenient, thereby improving the shooting experience of the user.
[0185] Exemplarily, the second interface can be Figures 3A-3C The user interface 33 in the embodiment can also be a user interface in other scenarios in specific implementation, which is not limited in the embodiment of the application.
[0186] Step S4B07: performing Fourier transform based on the first image to obtain a fifth image; the fifth image is a spectrum image of the first image converted from a spatial domain to a frequency domain;
[0187] Step S4B08: determining whether a specific frequency intensity greater than a preset value exists in the fifth image.
[0188] Step S4B09: if the specific frequency intensity greater than the preset value exists, it is determined that the first image includes the crosstalk artifact.
[0189] Specifically, for specific description of steps S4B07-S4B09, refer to the related description of step S4A05 above, which will not be repeated here.
[0190] Step S4B10: inputting the first image into a neural network model, correcting the first image through the neural network model to obtain a second image.
[0191] Specifically, the neural network model is used to remove the crosstalk artifact in the area with the crosstalk artifact in the first image; the second image is a RAW image after the crosstalk artifact is removed. For description of step S4B10, refer to the related description of step S4A06 above, which will not be repeated here.
[0192] Step S4B11: converting to obtain a third image based on the second image; the third image is an RGB image after the crosstalk artifact is removed.
[0193] Specifically, for specific description of step S4B11, refer to the related description of step S4A07 above, which will not be repeated here.
[0194] Step S4B12: receiving a third operation for the image preview frame.
[0195] Step S4B13: in response to the third operation, the display screen displays a third interface.
[0196] Specifically, the third interface includes a photo album display area, and the third image is displayed in the photo album display area. Exemplarily, the third operation can be Figure 3A In the embodiment, the operation of clicking the image preview frame 306, in response to the third operation, the third interface presented by the electronic device can be Figures 3A-3CThe user interface 34 provided in the embodiment allows the user to observe the RGB image (i.e., the third image) after removing crosstalk artifacts through the album display area 308 in the user interface 34. In specific implementations, the third operation may also be other operation methods, which are not limited in this embodiment.
[0197] Please see Figure 4C , Figure 4C This is a schematic flowchart illustrating how a user views an image after taking a picture, with artifacts removed, according to an embodiment of this application. Optionally, when viewing a photo after removing crosstalk artifacts, that is, in the above... Figure 4A Steps S4A01-S4A07, or Figure 4B Following steps S4B01-S4B13, the following may also be included: Figure 4C Steps S4C01-S4C05 in the process.
[0198] Step S4C01: Store the third image.
[0199] Step S4C02: Receive the fourth operation of launching the gallery application.
[0200] Step S4C03: In response to the fourth operation, the display screen shows the fourth interface.
[0201] The fourth interface includes a first display area, which is used to display thumbnails of stored images.
[0202] Step S4C04: Receive the fifth operation for the thumbnail of the third image.
[0203] Step S4C05: In response to the fifth operation, the display screen shows the fifth interface; the fifth interface includes a second display area for displaying the third image.
[0204] Specifically, the electronic device can store the RGB image (i.e., the third image) after removing crosstalk artifacts. When it receives a fourth operation to launch a gallery application, in response to the fourth operation, the display screen of the electronic device displays a fourth interface including a first display area for displaying thumbnails of the stored image. Further, when the electronic device receives a fifth operation for the thumbnail of the third image, in response to the fifth operation, the display screen of the electronic device displays a fifth interface including a second display area for displaying the third image. Through the embodiments of this application, users can view photos or videos after removing crosstalk artifacts through a gallery application, making it more convenient and faster for users to view images after removing crosstalk artifacts, thus improving the user's shooting experience.
[0205] For example, the fourth operation could be Figure 3CIn the embodiment, in the user interface 31, in response to the fourth operation, the fourth interface presented by the electronic device in response to the operation of clicking the gallery application can be Figures 3A-3C In the embodiment, the user interface 36 provided by the embodiment further includes that the fifth operation can be an operation of clicking the thumbnail 315 in the first display area 314, and at this time, the fifth interface can be the user interface 37 displayed by the display screen, and the second display area 316 in the user interface 37 can display the image 309. In a specific implementation, the above-mentioned operation mode (for example, the third operation and / or the fourth operation) can also be other operation modes, and the fourth interface and / or the fifth interface can also be other interfaces in other application scenarios, and the embodiment of the present application does not limit this.
[0206] In a possible implementation, the neural network model is trained based on a sample data set; the sample data set includes N pairs of data pairs, N is a positive integer; the data pair includes a sample image and a target image, the sample image is a RAW image with crosstalk artifacts, and the target image is a RAW image obtained by removing the crosstalk artifacts from a region with the crosstalk artifacts in the sample image based on a domain mean compensation method.
[0207] In a possible implementation, the step of removing the crosstalk artifacts from the region with the crosstalk artifacts in the sample image based on the domain mean compensation method can include: calculating the mean and variance of the channel value of each pixel in the pixel matrix of the same feature channel for the region with the crosstalk artifacts in the sample image; determining the adjusted target channel value by reducing the size of the variance and judging whether the intensity of the specific frequency is lower than a preset value; obtaining the target image with the crosstalk artifacts removed based on the target channel value.
[0208] Further, the crosstalk artifacts are removed from the region with the crosstalk artifacts in the sample image based on the domain mean compensation method. Please refer to Figure 5 , Figure 5 FIG. 1 is a schematic diagram of a domain mean compensation method for removing crosstalk artifacts provided by the embodiment of the present application.
[0209] As shown in Figure 5 , Figure 5 (A) in the above-mentioned (A) indicates the region with the crosstalk artifacts in the sample image, which includes a pixel matrix of multiple different feature channels (for example, R channel, G channel, B channel), and the channel value of each pixel is different. Since the pixel matrix of the sample image is obtained based on the multi-Bayer array of the image sensor, for the pixel matrix of the same feature channel in the region with the crosstalk artifacts in the sample image, it can be as shown in Figure 5The 2x2 matrix shown in (B) in the above formula can also be a pixel matrix of other arrangements (for example, a 3x3 matrix), and the embodiments of the present application do not limit the pixel matrix. Further, the mean and variance between each pixel of the pixel matrix of the same feature channel are calculated, for example Figure 5 The mean of the channel values (for example, R1, R2, R3, R4) of each pixel in the R matrix shown in (B) in the above formula is obtained, and the channel value of each pixel is R0; the variance between the channel value of each pixel and the average channel value R0 is calculated, and the variance matrix of the channel values R5, R6, R7, R8 is obtained. At this time, the original R matrix with channel values R1, R2, R3, R4 can be regarded as the sum of the mean matrix with channel R0 and the variance matrix with channel values R5, R6, R7, R8. Further, by reducing the variance, for example, multiplying the calculated variance of each pixel by a 0-1 coefficient to reduce the variance of each pixel, the channel value of each pixel of the pixel matrix with smaller luminance difference is calculated based on the reduced variance. For example, for the pixel matrix in the area with larger crosstalk artifacts, the coefficient multiplied by the pixel variance is smaller, that is, the amplitude of the variance reduction is larger; and for the pixel matrix in the area with smaller crosstalk artifacts, the coefficient multiplied by the pixel variance is smaller, that is, the amplitude of the variance reduction is smaller. Since the sensor can be more sensitive to a specific frequency of light intensity, which can cause crosstalk artifacts in the image, the present method reduces the variance to make the specific frequency of the pixel matrix after the domain mean compensation lower than the preset value, thereby removing the crosstalk artifacts. By the above method, the domain mean compensation is performed on all pixel matrices in the area with crosstalk artifacts in the sample image, and it is judged whether the intensity of the specific frequency of the obtained pixel matrix is lower than the preset value (usually 50HZ-60HZ), wherein the specific frequency is usually related to the refresh rate or reading frequency of the image sensor. If the specific frequency of the adjusted pixel matrix is lower than the preset value, the target image with crosstalk artifacts removed is obtained based on the adjusted target channel value. The target image removes crosstalk artifacts only in the area with crosstalk artifacts, so that the target image used for training the neural network can obtain the target image with crosstalk artifacts removed while not losing the image clarity to the greatest extent, so as to make the sample data set in this way to train the neural network model, so as to obtain the optimal neural network model for removing crosstalk artifacts, so as to improve the image authenticity and thus improve the user's shooting experience.
[0210] It should be noted that the above Figure 5 The above is only one possible implementation process of the domain mean compensation algorithm provided by the embodiments of the present application, and the specific implementation process can be changed based on different types of pixel matrices, and the embodiments of the present application do not limit this.
[0211] For example, please refer to Figure 6 , Figure 6A neural network model training diagram based on a U-Net architecture is provided for the embodiments of the present application. As an example, a sample image with crosstalk artifacts is input as an input picture for inputting the neural network model, and a target image after removing crosstalk artifacts in the region with crosstalk artifacts by the domain mean compensation algorithm described above is taken as the output picture of the neural network model. As shown in Figure 6 The U-net network shown in Figure has five layers in total, which performs 4 times of down-sampling and 4 times of up-sampling on the picture. For example, when the input picture is a single-channel picture of 572x572, it is then converted into a 64-channel feature map of 568x568 through two consecutive convolutions.
[0212] (1) The first red arrow in the upper left corner represents an operation that reduces the length and width of the feature map to half of the original. The size transformation here is from (568, 568, 64) to (284, 284, 64), and the subsequent two convolution layers increase the number of feature map channels to 128.
[0213] (2) The other operations in the left half are the same as the analysis method in (1).
[0214] (3) The number of feature map channels in the middle is 1024, and the subsequent up-sampling (or deconvolution) increases the length and width of the feature map.
[0215] (4) The gray arrow represents "copying" the feature map on the left to the feature map on the right in terms of channels. For example, the feature map with a size of (280, 280, 128) in the left half of the figure is connected to the feature map on the right (200, 200, 128) through the second gray arrow, resulting in a size of (200, 200, 256).
[0216] (5) Repeating the analysis in (4), it can be concluded that the final output image size is (388, 388, 64).
[0217] (6) Finally, through a 1x1 convolution, the number of output picture channels is reduced to 2.
[0218] Optionally, during the training process of the above-mentioned neural network model, in addition to the sample image with crosstalk artifacts and the target image after removing crosstalk artifacts as the input picture and the output picture of the neural network model, variance and other parameters can also be input to improve the speed and accuracy of model training, so as to achieve the best effect of model training.
[0219] Further based on the neural network model based on the U-Net architecture as shown in Figure 6 The training steps can include but are not limited to the following steps 1-4.
[0220] Step 1, obtaining a pair of data from a data set used for training the neural network model;
[0221] Step 2, inputting the sample image in the pair of data into the neural network model to obtain a corresponding sample corrected image;
[0222] Step 3, calculating a loss value between the target image and the corresponding sample corrected image according to a loss function;
[0223] Step 4, adjusting each parameter of the neural network model based on the calculated loss value, repeating steps 1-4 until the loss value reaches a preset index, and obtaining a trained neural network model.
[0224] Further, each pair of data includes a sample image with crosstalk artifacts and a target image obtained by removing the crosstalk artifacts in the region with crosstalk artifacts in the sample image based on a domain mean compensation method. By training the neural network model multiple times to adjust each parameter of the neural network model, the neural network model can only remove the crosstalk artifacts in the region with crosstalk artifacts in the sample image, and does not process the region without crosstalk artifacts, that is, the corrected image obtained by the neural network model based on the sample image is as close as possible to the target image corresponding to the sample image, thereby obtaining a trained neural network model that can be used to remove the crosstalk artifacts in the region with crosstalk artifacts in the first image, so as to remove the crosstalk artifacts by the neural network model to improve the image authenticity and thus improve the user's shooting experience.
[0225] It should be noted that the trained neural network model can be a neural network or a neural network based on other types of architecture; the loss function used to calculate the loss value between the target image and the corresponding sample corrected image can be an absolute value loss function (also known as L1 loss function), or other loss functions, which are not limited by the embodiments of the present application. The above steps 1-4 are only one possible implementation of training the neural network model, and do not constitute a specific limitation on the method for removing crosstalk artifacts provided by the embodiments of the present application.
[0226] The above detailed method of the embodiments of the present application can be understood that each device contains the corresponding hardware structure and / or software module for executing each function in order to realize the above corresponding functions. The units and steps of the examples described in conjunction with the embodiments disclosed herein can be implemented in hardware or a combination of hardware and computer software. Whether a function is executed by hardware or computer software driven hardware depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods for each specific application to implement the described functions, but such implementation should not be considered beyond the scope of the present application. The following describes the device provided by the embodiments of the present application.
[0227] The above describes the method of the embodiments of the application in detail. The device provided by the embodiments of the application is introduced below. Exemplarily, refer to Figure 7 , Figure 7 is a structural schematic diagram of a device for removing crosstalk artifacts provided by the embodiments of the application. The device 700 can be applied to an electronic device, which includes a display screen and a camera, the camera including an image sensor, wherein the pixels of the image sensor are arranged in a Bayer array arrangement. The device 700 can include a first receiving unit 701, a first display unit 702, a second receiving unit 703, a first obtaining unit 704, a first detecting unit 705, a first correcting unit 706, and a first converting unit 707. The detailed description of each unit is as follows:
[0228] The first receiving unit 701 is configured to receive a first operation of starting a camera application.
[0229] The first display unit 702 is configured to, in response to the first operation, display a first interface on the display screen. The first interface includes a viewfinder and a shooting control. The viewfinder is configured to display a first picture in real time. The first picture is an RGB image of a current shooting scene converted from a RAW image collected by the image sensor in real time.
[0230] The second receiving unit 703 is configured to receive a second operation acting on the shooting control.
[0231] The first obtaining unit 704 is configured to, in response to the second operation, obtain a first image. The first image is a RAW image corresponding to the first picture collected by the image sensor after the second operation is received.
[0232] The first detecting unit 705 is configured to detect whether the first image includes crosstalk artifacts.
[0233] The first correcting unit 706 is configured to, if the first image includes crosstalk artifacts, correct the first image based on a preset neural network model to obtain a second image. The neural network model is configured to remove crosstalk artifacts in a crosstalk artifact region of the first image. The second image is a RAW image after crosstalk artifacts are removed.
[0234] The first converting unit 707 is configured to convert the second image to obtain a third image. The third image is an RGB image after crosstalk artifacts are removed.
[0235] The device for removing crosstalk artifacts provided by the embodiment of the present application is applied to an electronic device. After the first receiving unit 701 receives a first operation of starting a camera application, the first display unit 702 responds to the first operation to display, through the display screen of the electronic device, a first interface including a viewfinder frame and a shooting control. The viewfinder frame is used to display, in real time, an RGB image of a current shooting scene (that is, a first picture) converted from a RAW image collected by the image sensor in real time. Further, after the second receiving unit 703 receives a second operation acting on the shooting control, the first obtaining unit 704 responds to the second operation to obtain a RAW image (that is, a first image) corresponding to the first picture collected by the image sensor after the second operation is received. Further, the first detecting unit 705 detects whether the first image includes crosstalk artifacts. If it is detected that the first image includes crosstalk artifacts, the first correcting unit 706 corrects the first image based on a preset neural network model to remove crosstalk artifacts in the region with crosstalk artifacts in the first image, to obtain a RAW image after crosstalk artifacts are removed (that is, a second image), and then the first converting unit 707 converts the second image to obtain an RGB image after crosstalk artifacts are removed (that is, a third image). Since the pixel arrangement of the image sensor is a multi-Bayer array, that is, a plurality of pixels of the same color are arranged on the original pixel, the spacing between adjacent pixels is small, and the crosstalk problem between pixels of the same color is aggravated, so that the RGB image converted by the electronic device based on the RAW image (for example, the first image) is prone to crosstalk artifacts, such as grid-shaped crosstalk artifacts, thereby reducing the authenticity of the captured image. Therefore, the embodiment of the present application can first detect the RAW image corresponding to the first picture collected by the image sensor after the electronic device receives the second operation, that is, by detecting the first image, whether the first image includes crosstalk artifacts. Further, when it is detected that the RAW image (that is, the first image) obtained by the electronic device based on the image sensor includes crosstalk artifacts, the preset neural network model is used to remove only the crosstalk artifacts in the region with crosstalk artifacts in the first image, and the region without crosstalk artifacts in the first image is not processed, to obtain the RAW image after crosstalk artifacts are removed (that is, the second image), so that the RGB image after crosstalk artifacts are removed (that is, the third image) converted based on the RAW image has an improved region with crosstalk artifacts, and the region without crosstalk artifacts can continue to maintain the previous clarity, thereby greatly restoring the authenticity and clarity of the image, and improving the user's shooting experience.
[0236] In a possible implementation, the device further includes:
[0237] The second acquisition unit is configured to acquire a fourth image after receiving a second operation acting on the shooting control; the fourth image is an RGB image corresponding to the first image after receiving the second operation;
[0238] The second display unit is configured to display a second interface on the display screen; the second interface includes an image preview frame, and the fourth image is displayed in the image preview frame;
[0239] The third receiving unit is configured to receive a third operation on the image preview frame;
[0240] The third display unit is configured to display a third interface on the display screen in response to the third operation; the third interface includes a photo album display area, and the third image is displayed in the photo album display area.
[0241] In a possible implementation, the second interface further includes a preview photo saving control, and the third operation includes an operation acting on the preview photo saving control; or the third operation includes a viewing operation on the image preview frame.
[0242] In a possible implementation, the apparatus further includes:
[0243] The first storage unit is configured to store the third image;
[0244] The fourth receiving unit is configured to receive a fourth operation of starting a photo album application;
[0245] The fourth display unit is configured to display a fourth interface on the display screen in response to the fourth operation; the fourth interface includes a first display area, and the first display area is configured to display thumbnails of stored images;
[0246] The fifth receiving unit is configured to receive a fifth operation on the thumbnail of the third image;
[0247] The fifth display unit is configured to display a fifth interface on the display screen in response to the fifth operation; the fifth interface includes a second display area, and the second display area is configured to display the third image.
[0248] In a possible implementation, the first detection unit 705 is specifically configured to:
[0249] perform Fourier transform based on the first image to obtain a fifth image; the fifth image is a frequency spectrum image converted from a spatial domain to a frequency domain;
[0250] determine whether a specific frequency intensity in the fifth image is greater than a preset value;
[0251] If there is a specific frequency intensity greater than the preset value, it is determined that the first image includes crosstalk artifacts.
[0252] In a possible implementation, the first correction unit 706 is specifically configured to:
[0253] input the first image into the neural network model, correct the first image through the neural network model, and obtain the second image.
[0254] In a possible implementation, the neural network model is obtained based on a sample data set; the sample data set includes N pairs of data, N being a positive integer; the data pair includes a sample image and a target image, the sample image being a RAW image with crosstalk artifacts, and the target image being a RAW image obtained by removing the crosstalk artifacts from a region with the crosstalk artifacts in the sample image based on a domain mean compensation method.
[0255] In a possible implementation, the removing the crosstalk artifacts from the region with the crosstalk artifacts in the sample image based on the domain mean compensation method includes:
[0256] calculating a mean and a variance of channel values of each pixel in a pixel matrix of the same feature channel for the region with the crosstalk artifacts in the sample image;
[0257] determining an adjusted target channel value by reducing the size of the variance and judging whether the intensity of the specific frequency is lower than a preset value;
[0258] obtaining the target image with the crosstalk artifacts removed based on the target channel value.
[0259] It should be noted that the functions of each unit in the apparatus 700 described in the embodiments of the present application can refer to the related descriptions of the above method embodiments, which will not be described here. It can be understood that the apparatus and method provided by the embodiments of the present application can be implemented in other ways. For example, the above-described system embodiments are only schematic, for example, the division of the above-mentioned modules or units is only a logical function division, and actual implementation can have another division manner, for example, a plurality of units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the coupling or direct coupling or communication connection between the displayed or discussed each other can be through some interface, indirect coupling or communication connection between the units or components, which can be electrical, mechanical or other forms.
[0260] Exemplarily, please refer to Figure 8 , Figure 8 is another hardware structure schematic diagram of an electronic device provided by the embodiments of the present application. As shown in FIG. 8, the electronic device 800 includes a processor 801, a memory 802 and a communication interface 803. The processor 801, the memory 802 and the communication interface 803 are connected through a bus. The bus can be a hardware component line, a wireless bus or a combination of the two. The bus can be a wired bus or a wireless bus. The bus can be a combination of a wired bus and a wireless bus.Figure 8 As shown, the electronic device 800 includes at least one processor 801 and a memory 802. The processor 801 is coupled with the memory 802, and the coupling in the embodiments of the present application can be a communication connection, can be electrical, or other forms. In addition, the electronic device 800 provided by the embodiments of the present application can further include a camera 803 and at least one display screen 804 (only one is shown), the camera including an image sensor 805. The processor 801, the memory 802, the camera 803, and the display screen 804 can be connected through a bus 806. Specifically, the memory 802 is configured to store program instructions. The processor 801 is configured to invoke the program instructions stored in the memory 802, so that the electronic device 800 can execute the steps in the control method provided by the embodiments of the present application. The description of each component and related steps can be referred to the above, and will not be repeated here. Figure 8 It should be noted that the electronic device 800 provided by the embodiments of the present application can include more or fewer components than shown, or combine certain components, or split certain components, or different component arrangements. The components shown can be implemented in hardware, software, or any combination of software and hardware.
[0261] It should be noted that the electronic device 800 provided by the embodiments of the present application can include more or fewer components than shown, or combine certain components, or split certain components, or different component arrangements. The components shown can be implemented in hardware, software, or any combination of software and hardware.
[0262] The embodiments of the present application provide a computer readable storage medium, which stores a computer program, the computer program is executed by the processor of the above-mentioned routing device to realize the steps executed by the routing device in the above-mentioned method for recovering the network provided by the embodiments of the present application; or the computer program is executed by the processor of the above-mentioned electronic device to realize the steps executed by the electronic device in the above-mentioned method for recovering the network provided by the embodiments of the present application.
[0263] The embodiments of the present application provide a computer program, which includes instructions, the computer program is executed by the processor of the above-mentioned routing device to realize the steps executed by the routing device in the above-mentioned method for recovering the network provided by the embodiments of the present application; or the computer program is executed by the processor of the above-mentioned electronic device to realize the steps executed by the electronic device in the above-mentioned method for recovering the network provided by the embodiments of the present application.
[0264] In the above embodiments, the description of each embodiment has its own emphasis, and the parts not described in detail in a certain embodiment can be referred to the related description of other embodiments.
[0265] It should be noted that, for the aforementioned method embodiments, the purposes of simple description, they are all expressed as a series of action combinations, but those skilled in the art should know that the present application is not limited to the action sequence described, because according to the present application, certain steps can be performed in other sequences or simultaneously, or certain steps can be omitted, a plurality of steps can be combined into one step, and / or one step can be divided into a plurality of steps. Secondly, those skilled in the art should know that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily required by the present application. It should be noted that the features and functions of two or more devices according to the present disclosure can be embodied in one device. Conversely, the features and functions of one device described above can be further divided into devices.
[0266] In the above embodiments, all or part can be realized by software, hardware, firmware or any combination thereof. When realized by software, all or part can be realized in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, all or part of the processes or functions described in the present application are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer readable storage medium or transferred from one computer readable storage medium to another, for example, the computer instructions can be transferred from one website, computer, server or data center to another website, computer, server or data center through wired (such as coaxial cable, optical fiber, digital subscriber line) or wireless (such as infrared, wireless, microwave, etc.) mode. The computer readable storage medium can be any available medium that a computer can access or a data storage device such as a server, data center, etc. that includes one or more available media sets. The available media can be magnetic media (for example, floppy disk, hard disk, magnetic tape), optical media (for example, DVD), or semiconductor media (for example, solid state disk (SSD)) and the like.
[0267] Those of ordinary skill in the art can understand that all or part of the processes in the above-mentioned embodiments can be instructed by a computer program to complete the relevant hardware, which can be stored in a computer readable storage medium. The program can include the processes of the above-mentioned method embodiments when executed. The aforementioned storage medium includes ROM or random storage memory RAM, magnetic disk or optical disk and various program code storage media.
[0268] In conclusion, the above-mentioned is only the embodiment of the technical scheme of the present application, and is not used to limit the protection scope of the present application. Any modification, equivalent replacement, improvement, etc. made according to the disclosure of the present application shall be included in the protection scope of the present application.
Claims
1. A method of removing crosstalk artifacts, characterized by, The method is applied to an electronic device, the electronic device comprising a display screen and a camera, the camera comprising an image sensor, wherein pixels of the image sensor are arranged in a Bayer array; the method comprising: receiving a first operation for starting a camera application; in response to the first operation, the display screen displays a first interface; wherein the first interface comprises a viewfinder and a shooting control, the viewfinder is used to display a first picture in real time, and the first picture is an RGB image of a current shooting scene converted from a RAW image collected by the image sensor in real time; receiving a second operation acting on the shooting control; in response to the second operation, a first image is obtained; the first image is a RAW image corresponding to the first picture collected by the image sensor after the second operation is received; detecting whether the first image comprises crosstalk artifacts; if the first image comprises crosstalk artifacts, correcting the first image based on a preset neural network model to obtain a second image; wherein the neural network model is used to remove crosstalk artifacts in a region with crosstalk artifacts in the first image, and is not used to process a region without crosstalk artifacts in the first image; the second image is a RAW image after crosstalk artifacts are removed; based on the second image, a third image is converted; the third image is an RGB image after crosstalk artifacts are removed; wherein the neural network model is obtained based on a sample data set; wherein the sample data set comprises N pairs of data, N being a positive integer; the data pair comprises a sample image and a target image, the sample image is a RAW image with crosstalk artifacts, and the target image is a RAW image obtained by removing crosstalk artifacts in a region with crosstalk artifacts in the sample image based on a domain mean compensation method; the method for removing crosstalk artifacts in the region with crosstalk artifacts in the sample image based on the domain mean compensation method comprises: calculating the mean and variance of channel values of each pixel in a pixel matrix of the same feature channel for the region with crosstalk artifacts in the sample image; determining an adjusted target channel value by reducing the size of the variance and judging whether the intensity of a specific frequency is lower than a preset value; obtaining the target image with crosstalk artifacts removed based on the target channel value.
2. The method of claim 1, wherein, The method further comprises: after the second operation acting on the shooting control is received, a fourth image is obtained; the fourth image is an RGB image corresponding to the first picture after the second operation is received; the display screen displays a second interface; the second interface comprises an image preview frame, and the fourth image is displayed in the image preview frame; receiving a third operation for the image preview frame; in response to the third operation, the display screen displays a third interface; the third interface comprises an album display area, and the third image is displayed in the album display area.
3. The method of claim 2, wherein, The second interface further includes a preview photo saving control, and the third operation includes an operation on the preview photo saving control; or the third operation includes a viewing operation on the image preview frame.
4. The method according to any one of claims 1 to 3, characterized in that, The method further includes: storing the third image; receiving a fourth operation of starting a gallery application; in response to the fourth operation, the display screen displays a fourth interface; the fourth interface includes a first display area for displaying thumbnails of stored images; receiving a fifth operation on the thumbnail of the third image; in response to the fifth operation, the display screen displays a fifth interface; the fifth interface includes a second display area for displaying the third image.
5. The method according to any one of claims 1 to 3, characterized in that, The detection of whether the first image includes a crosstalk artifact includes: performing Fourier transform on the first image to obtain a fifth image; the fifth image is a frequency spectrum image converted from a spatial domain to a frequency domain; determining whether a specific frequency intensity in the fifth image is greater than a preset value; if the specific frequency intensity is greater than the preset value, it is determined that the first image includes a crosstalk artifact.
6. The method according to any one of claims 1 to 3, characterized in that, The correction of the first image based on a preset neural network model to obtain a second image includes: inputting the first image into the neural network model to correct the first image through the neural network model to obtain the second image.
7. An apparatus for removing crosstalk artifacts, the apparatus comprising: An electronic device includes a display screen and a camera, and the camera includes an image sensor, where the pixels of the image sensor are arranged in a multi-Bayer array; the electronic device includes: a first receiving unit configured to receive a first operation of starting a camera application; a first display unit configured to, in response to the first operation, display a first interface on the display screen; the first interface includes a viewfinder frame and a shooting control, and the viewfinder frame is configured to display a first picture, which is an RGB image of a current shooting scene converted from a RAW image collected by the image sensor in real time; a second receiving unit configured to receive a second operation on the shooting control; a first obtaining unit configured to, in response to the second operation, obtain a first image; the first image is a RAW image corresponding to the first picture collected by the image sensor after the second operation is received; a first detection unit configured to detect whether the first image includes a crosstalk artifact; a first correction unit configured to, if the first image includes a crosstalk artifact, correct the first image based on a preset neural network model to obtain a second image; the neural network model is configured to remove the crosstalk artifact in the region with the crosstalk artifact in the first image; the second image is a RAW image after the crosstalk artifact is removed; a first conversion unit configured to convert the second image to obtain a third image; the third image is an RGB image after the crosstalk artifact is removed. The neural network model is obtained based on a sample data set; the sample data set includes N pairs of data pairs, N being a positive integer; the data pairs include a sample image and a target image, the sample image being a RAW image with crosstalk artifacts, and the target image being a RAW image obtained by removing the crosstalk artifacts from a region with the crosstalk artifacts in the sample image based on a domain mean compensation method; the removing the crosstalk artifacts from the region with the crosstalk artifacts in the sample image based on the domain mean compensation method includes: calculating a mean value and a variance of channel values of each pixel in a pixel matrix of a same feature channel for the region with the crosstalk artifacts in the sample image; determining an adjusted target channel value by reducing the size of the variance and judging whether the intensity of a specific frequency is lower than a preset value; and obtaining the target image with the crosstalk artifacts removed based on the target channel value.
8. An electronic device, comprising: The electronic device includes a display screen and a camera, the camera including an image sensor, wherein the pixels of the image sensor are arranged in a Bayer array; the electronic device further includes a memory and a processor, wherein the memory is configured to store a computer program, and the processor is configured to execute the computer program stored in the memory, so that the electronic device executes the method of any one of claims 1-6.
9. A computer storable medium, characterized by The computer readable storage medium stores a computer program, and the computer program is executed by the processor to implement the method of any one of claims 1-6.
10. A computer program product, characterised in that, The computer program product includes instructions executed by a computing device to implement the method of any one of claims 1-6. The computer program product includes instructions executed by a computing device to implement the method of any one of claims 1-6.
Citation Information
Patent Citations
Image display method and device
CN111741211A
Method and apparatus for processing image artifact by using electronic device
WO2021141216A1