Image processing method and related device

By fusing and interpolating multiple frames captured under a stroboscopic light source, or utilizing the difference between long and short frames of a staggered HDR camera, the stripes are effectively removed, improving image quality.

CN120751271AActive Publication Date: 2025-10-03HONOR DEVICE CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202411218061.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-08-30
Publication Date
2025-10-03
Estimated Expiration
2044-08-30

AI Technical Summary

Technical Problem

In images captured under a stroboscopic light source, striping phenomena (such as bright stripes and dark stripes) are easily present, which are difficult to effectively remove using existing technologies.

Method used

A fused image is obtained by performing a mean operation on multiple frame images, and the difference between the fused image and the image to be processed is calculated to obtain a stripe mapping image. Stripes are then removed based on a weighted fusion method, or a stagger HDR camera is used to obtain a difference image between long and short frames for stripe removal.

Benefits of technology

It effectively removes banding from images and improves image quality, especially for images of static or moving objects captured under stroboscopic light sources.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120751271A_ABST
    Figure CN120751271A_ABST
Patent Text Reader

Abstract

The invention provides an image processing method and a related device, which are applied to electronic equipment, a camera running in the electronic equipment obtains multiple frames of images by calling a camera to shoot an object with a first exposure time length, the object is in an environment with a stroboscopic light source and is static, and the first exposure time length is smaller than a stroboscopic period. The method comprises the steps of performing mean value operation on multiple frames of images to obtain a fused image, obtaining a stripe mapping image based on the difference between the fused image and an image to be processed, fusing the fused image and the image to be processed based on the stripe mapping image, and obtaining a stripe-removed image, and in the fusion process, the larger the value of a first pixel point in the stripe mapping image is, the larger the value of a second pixel point in the stripe mapping image is. The weight of the fused image at the first pixel point is larger, and the weight of the to-be-processed image at the first pixel point is smaller, so that the influence of the stripe-free fused image on the fused first pixel point after fusion is larger, and the purpose of removing stripes is achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of digital signal processing technology, and in particular to an image processing method and related devices. Background Art

[0002] A stroboscopic light source can be understood as a light source powered by alternating current with a certain frequency, such as a common fluorescent lamp.

[0003] Images taken under stroboscopic light may have banding, also known as water ripples, etc. Figure 1 For example, stripes appear in photos taken indoors with the lights on, including bright stripes and dark stripes. Summary of the Invention

[0004] This application provides an image processing method and related devices to remove stripes from images. The disclosed technical solutions are as follows:

[0005] A first aspect of the present application provides an image processing method, which is applied to an electronic device, wherein a camera running in the electronic device obtains multiple frames of images by calling the camera to shoot an object with a first exposure time, the object is in an environment with a stroboscopic light source and is stationary, and the first exposure time is less than the period of the stroboscopic light, and the method includes: performing a mean operation on the multiple frames of images to obtain a fused image, obtaining a strip mapping image based on the difference between the fused image and the image to be processed, and fusing the fused image with the image to be processed based on the strip mapping image to obtain an image without stripes, and during the fusion process, obtaining a value of a first pixel point in the image without stripes based on a weighted fusion method of the value of the first pixel point in the fused image and the value of the first pixel point in the image to be processed, wherein the larger the value of the first pixel point in the strip mapping image, the greater the weight of the fused image at the first pixel point, and the smaller the weight of the image to be processed at the first pixel point.

[0006] Based on the characteristics of the camera and the environment, it can be seen that both the multi-frame images and the image to be processed contain strips, but the positions of the strips in the multi-frame images are different. Therefore, the fused image obtained by the mean operation does not contain stripes. Therefore, in the stripe mapping image obtained based on the difference between the fused image and the image to be processed, the larger the value of the pixel point, the greater the difference between the fused image and the image to be processed at the pixel point, and therefore the more likely the pixel point is a pixel point on the stripe. Therefore, in the fusion process, the larger the value of the first pixel point (which can be any pixel point) in the stripe mapping image, the greater the weight of the fused image at the first pixel point, and the smaller the weight of the image to be processed at the first pixel point, so that the pixel point after fusion is more affected by the stripe-free fused image, thereby achieving the purpose of removing stripes.

[0007] In some implementations, obtaining a stripe mapping image based on the difference between the fused image and the image to be processed includes: obtaining the difference between the fused image and the image to be processed to obtain a stripe mask image; and extracting a dark stripe mask image and a bright stripe mask image from the stripe mask image to obtain the stripe mapping image. Extracting the dark stripe mask image and the bright stripe mask image facilitates separate processing of dark strips and bright strips, thereby improving the accuracy of the stripe mapping image and further enhancing the effectiveness of stripe removal.

[0008] In some implementations, extracting a dark stripe mask image and a bright stripe mask image from a stripe mask image includes: retaining pixel values ​​in the stripe mask image that are greater than or equal to a first threshold (such as threshold value thr_1) or setting them to 1, and setting pixel values ​​that are less than the first threshold to 0, thereby obtaining a dark stripe mask image; retaining pixel values ​​in a flipped image of the stripe mask image that are greater than or equal to the first threshold or setting them to 1, and setting pixel values ​​that are less than the first threshold to 0, thereby obtaining a bright stripe mask image.

[0009] A second aspect of the present application provides an image processing method, which is applied to an electronic device. A camera running in the electronic device captures an object by calling a stacked high dynamic range (stagger HDR) camera. The object is in an environment with a stroboscopic light source. The exposure time of the long frame captured by the stagger HDR camera is greater than or equal to the stroboscopic period, and the exposure time of the short frame captured by the stagger HDR camera is less than the stroboscopic period. The method includes: obtaining a strip mapping image based on the difference between a first long frame and a first short frame. The first long frame and the first short frame have the same shooting time. Based on characteristics such as the exposure time of the stagger HDR camera and characteristics of the ambient light source, the difference between the first long frame and the second long frame is a difference caused by stripes, and does not include differences caused by motion, that is, the strip mapping image can reflect such differences.

[0010] Based on the stripe-mapped image, a first image and a second image are fused to obtain a stripe-free image. The first image is obtained based on the first long frame, and the second image is obtained based on the first short frame. During the fusion process, the value of a first pixel in the stripe-mapped image is obtained by weighted fusion of the values ​​of the first pixel in the first image and the values ​​of the first pixel in the second image. A larger value of the first pixel in the stripe-mapped image indicates a greater weight of the first image at the first pixel, and a smaller weight of the second image at the first pixel. The first image can be the first long frame. To obtain a higher-quality stripe-free image, the first image can also be an image obtained by registering the first long frame with the first short frame and aligning the brightness. Similarly, the second image can be the first short frame or an image obtained by registering the first short frame with the first long frame. A larger value of the first pixel in the stripe-mapped image indicates a greater weight of the first image at the first pixel, and a smaller weight of the second image at the first pixel. This results in a greater influence of the stripe-free fused image on the first pixel after fusion, thereby achieving stripe removal.

[0011] In some implementations, obtaining a stripe mapping image based on the difference between the first long frame and the first short frame includes: calculating the difference between the first long frame and the first short frame to obtain a difference mask image, the difference mask image representing the difference between the first long frame and the first short frame caused by object motion and stripes; obtaining a stripe mask image based on a previously obtained motion mask image and a difference mask image, the motion mask image representing the difference between the first long frame and an adjacent long frame caused by object motion; the stripe mask image representing the difference between the first long frame and the first short frame caused by stripes; and extracting a dark stripe mask image and a bright stripe mask image from the stripe mask image to obtain a stripe mapping image. Removing the difference caused by motion from the difference mask image facilitates obtaining a stripe mapping image that more accurately represents stripes, thereby improving the accuracy of stripe removal.

[0012] In some implementations, before obtaining the strip mask image based on the pre-acquired motion mask image and the difference mask image, the method further includes: obtaining a motion mask image based on the difference between the first long frame and the second long frame, wherein the motion mask image represents the difference between the first long frame and the second long frame caused by the motion of the object, and the first long frame and the second long frame are image frames with adjacent timestamps, so as to obtain more accurate motion differences.

[0013] In some implementations, extracting a dark stripe mask image and a bright stripe mask image from a stripe mask image includes: retaining pixel values ​​in the stripe mask image that are greater than or equal to a first threshold (such as threshold value thr_2) or setting them to 1, and setting pixel values ​​less than the first threshold to 0, to obtain a dark stripe mask image; retaining pixel values ​​in a negated image of the stripe mask image that are greater than or equal to the first threshold or setting them to 1, and setting pixel values ​​less than the first threshold to 0, to obtain a bright stripe mask image.

[0014] In some implementations, before fusing the first image and the second image based on the strip mapping image, the process further includes: registering the first long frame and the first short frame to obtain a registered long frame and a registered short frame, performing brightness alignment on the registered long frame to the registered short frame to obtain a brightness-processed long frame, the first image being the brightness-processed long frame, and the second image being the registered short frame, aligning the positions and brightness of the first long frame and the first short frame, which is beneficial to improving the quality of the strip-processed image.

[0015] The third aspect of the present application provides an electronic device, comprising: one or more processors, a memory and a touch screen, the memory being used to store program code, and the processor being used to run the program code, so that the electronic device implements the image processing method provided in the first aspect or the second aspect of the present application.

[0016] The fourth aspect of the present application provides a computer-readable storage medium having instructions stored thereon. When the instructions are executed on an electronic device, the electronic device executes the image processing method provided in the first aspect or the second aspect of the present application. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0018] Figure 1 is an example image containing stripes;

[0019] Figure 2 An example diagram of the structure of an electronic device provided in an embodiment of the present application;

[0020] Figure 3 is an example diagram of a software framework running in an electronic device provided in an embodiment of the present application;

[0021] Figure 4 This is a flowchart of an image processing method provided by an embodiment of the present application;

[0022] Figure 5 This is a flowchart of another image processing method provided by an embodiment of the present application;

[0023] Figure 6 This is a flowchart of obtaining a BandingMap in the image processing method provided in an embodiment of the present application;

[0024] Figure 7 This is a flowchart of another image processing method provided by an embodiment of the present application;

[0025] Figure 8 This is an example diagram of obtaining weights based on BandingMap in the image processing method provided in the embodiment of the present application;

[0026] Figure 9 This is a flowchart of another image processing method provided by an embodiment of the present application;

[0027] Figure 10 This is a flowchart of another image processing method provided in an embodiment of the present application. DETAILED DESCRIPTION

[0028] The terms "first", "second" and "third" in the specification, claims and drawings of this application are used to distinguish different objects rather than to limit a specific order.

[0029] In the embodiments of the present application, words such as "in some implementations" or "for example" are used to indicate examples, illustrations or explanations, and should not be interpreted as being more preferred or more advantageous than other embodiments or design solutions.

[0030] In order to remove stripes in an image, an embodiment of the present application discloses an image processing method, which is applied to an electronic device.

[0031] Electronic devices include but are not limited to mobile phones, tablet computers, desktop computers, laptop computers, notebook computers, ultra-mobile personal computers (UMPCs), handheld computers, netbooks, personal digital assistants (PDAs), tablet computers (PADs), and wearable electronic devices such as smart watches and other electronic devices with cameras.

[0032] like Figure 2 As shown, taking a mobile phone as an example, the electronic device 100 may include a processor 110, an internal memory 120, a display screen 130, a camera 140, an antenna 1, an antenna 2, a mobile communication module 150, a wireless communication module 160 and an audio module 170, etc.

[0033] It should be understood that the structure illustrated in this embodiment does not constitute a specific limitation on the electronic device. In other embodiments, the electronic device may include more or fewer components than shown, or some components may be combined or separated, or the components may be arranged differently. The components shown in the figure may be implemented in hardware, software, or a combination of software and hardware.

[0034] The processor 110 may include one or more processing units. For example, the processor 110 may include an application processor (AP), a modem processor, a graphics processing unit (GPU), an image signal processor (ISP), a controller, a video codec, a digital signal processor (DSP), a baseband processor, and / or a neural-network processing unit (NPU). The different processing units may be independent devices or integrated into one or more processors.

[0035] The internal memory 120 can be used to store computer executable program codes, and the executable program codes include instructions. The processor 110 executes various functional applications and data processing of the electronic device 100 by running the instructions stored in the internal memory 110. The internal memory 120 may include a program storage area and a data storage area. Among them, the program storage area may store an operating system, an application required for at least one function (such as a sound playback function, an image playback function, etc.), etc. The data storage area may store data created during the use of the electronic device (such as audio data, a phone book, etc.), etc. In addition, the internal memory 120 may include a high-speed random access memory, and may also include a non-volatile memory, such as at least one disk storage device, a flash memory device, a universal flash storage (UFS), etc. The processor 110 executes various functional applications and data processing of the electronic device by running the instructions stored in the internal memory 120, and / or the instructions stored in the memory provided in the processor.

[0036] The electronic device implements its display function through a GPU, display screen 130, and an application processor. The GPU is a microprocessor for image processing that connects the display screen 130 and the application processor. The GPU is used to perform mathematical and geometric calculations for graphics rendering. The processor 110 may include one or more GPUs, which execute program instructions to generate or modify display information.

[0037] The display screen 130 is used to display images, videos, and the like. The display screen 130 includes a display panel. The display panel can be a liquid crystal display (LCD), an organic light-emitting diode (OLED), an active-matrix organic light-emitting diode (AMOLED), a flexible light-emitting diode (FLED), a MiniLED, a MicroLED, a Micro-OLED, or a quantum dot light-emitting diode (QLED). In some embodiments, the electronic device can include one or N display screens 130, where N is a positive integer greater than one.

[0038] The electronic device 100 can implement a shooting function through an ISP, a camera 140, a video codec, a GPU, a display screen 194, and an application processor.

[0039] The ISP processes data fed back by camera 140. For example, when taking a photo, the shutter is opened, and light is transmitted through the lens to the camera's photosensitive element. The light signal is converted into an electrical signal, which is then passed to the ISP for processing and transformed into a visible image. The ISP can also perform algorithmic optimization on image noise, brightness, and skin tone. It can also optimize parameters such as exposure and color temperature of the captured scene. In some embodiments, the ISP can be located within camera 140.

[0040] The camera 140 is used to capture still images or videos. The object generates an optical image through the lens and projects it onto the photosensitive element. The photosensitive element can be a charge coupled device (CCD) or a complementary metal-oxide-semiconductor (CMOS) phototransistor. The photosensitive element converts the light signal into an electrical signal, and then passes the electrical signal to the ISP for conversion into a digital image signal. The ISP outputs the digital image signal to the DSP for processing. The DSP converts the digital image signal into an image signal in a standard RGB, YUV or other format. In some embodiments, the electronic device 100 may include 1 or N cameras 140, where N is a positive integer greater than 1.

[0041] In some embodiments, the camera 140 is a camera with an exposure time shorter than the stroboscopic period of the ambient light source. In other embodiments, the camera 140 is a staggered high-dynamic range (staggerHDR) camera. The characteristics of the camera 140 result in stripes in the captured image, such as Figure 1 As shown, it will be described in detail in conjunction with subsequent embodiments.

[0042] The digital signal processor is used to process digital signals. In addition to processing digital image signals, it can also process other digital signals. For example, when the electronic device 100 selects a frequency point, the digital signal processor is used to perform Fourier transform on the frequency point energy.

[0043] Video codecs are used to compress or decompress digital video. Electronic device 100 may support one or more video codecs. This allows electronic device 100 to play or record videos in various encoding formats, such as Moving Picture Experts Group (MPEG) 1, MPEG2, MPEG3, and MPEG4.

[0044] The internal memory 120 can be used to store computer executable program codes, which include instructions. The processor 110 executes various functional applications and data processing of the electronic device 100 by running the instructions stored in the internal memory 120. The internal memory 120 may include a program storage area and a data storage area. Among them, the program storage area may store an operating system, an application required for at least one function (such as a sound playback function, an image playback function, etc.), etc. The data storage area may store data created during the use of the electronic device 100 (such as audio data, a phone book, etc.), etc. In addition, the internal memory 120 may include a high-speed random access memory, and may also include a non-volatile memory, such as at least one disk storage device, a flash memory device, a universal flash storage (UFS), etc. The processor 110 executes various functional applications and data processing of the electronic device 100 by running the instructions stored in the internal memory 120, and / or the instructions stored in the memory provided in the processor.

[0045] The electronic device 100 can implement audio functions such as music playback and recording through the audio module 170, the speaker 170A, the microphone 170B, and the application processor.

[0046] The audio module 170 is used to convert digital audio information into analog audio signal output, and is also used to convert analog audio input into digital audio signals. The audio module 170 can also be used to encode and decode audio signals. In some embodiments, the audio module 170 can be provided in the processor 110, or some functional modules of the audio module 170 can be provided in the processor 110.

[0047] The speaker 170A, also called a "speaker", is used to convert audio electrical signals into sound signals. The electronic device 100 can listen to music or listen to hands-free calls through the speaker 170A.

[0048] In some embodiments, the speaker 170A can play the video information with special effects mentioned in the embodiments of the present application.

[0049] Microphone 170B, also known as a "microphone" or "speaker", is used to convert sound signals into electrical signals. When making a call or sending a voice message, the user can speak by putting their mouth close to microphone 170B to input the sound signal into microphone 170B.

[0050] In some embodiments, the microphone 170B can collect sounds of the environment in which the electronic device is located while the camera is shooting video information with special effects.

[0051] The wireless communication function of the electronic device 100 can be implemented through the antenna 1, the antenna 2, the mobile communication module 150, the wireless communication module 160, the modem processor and the baseband processor.

[0052] Antenna 1 and antenna 2 are used to transmit and receive electromagnetic wave signals.

[0053] The mobile communication module 150 can provide wireless communication solutions, including 2G / 3G / 4G / 5G, for the electronic device 100. The mobile communication module 150 may include at least one filter, a switch, a power amplifier, a low-noise amplifier (LNA), and the like. The mobile communication module 150 can receive electromagnetic waves from the antenna 1, filter and amplify the received electromagnetic waves, and transmit them to the modem processor for demodulation. The mobile communication module 150 can also amplify the signals modulated by the modem processor and convert them into electromagnetic waves for radiation via the antenna 1.

[0054] The wireless communication module 160 can provide wireless communication solutions including wireless local area networks (WLAN) (such as wireless fidelity (Wi-Fi) networks), Bluetooth (BT), global navigation satellite system (GNSS), frequency modulation (FM), near field communication (NFC), infrared (IR), etc., which are applied to the electronic device 100. The wireless communication module 160 can be one or more devices integrating at least one communication processing module. The wireless communication module 160 receives electromagnetic waves via the antenna 2, frequency modulates and filters the electromagnetic wave signals, and sends the processed signals to the processor 110. The wireless communication module 160 can also receive the signal to be sent from the processor 110, frequency modulate it, amplify it, and convert it into electromagnetic waves for radiation through the antenna 2.

[0055] An operating system runs on top of the above components, such as the iOS operating system, Android operating system, and Windows operating system. Applications can be installed and run on the operating system.

[0056] Figure 3 It is a software structure block diagram of the electronic device according to an embodiment of the present application.

[0057] A layered architecture divides software into several layers, each with distinct roles and responsibilities. Layers communicate with each other through software interfaces. In some embodiments, the Android system is divided into five layers: from top to bottom: the application layer, the application framework layer, the Android runtime and system libraries, the hardware abstraction layer (HAL), and the kernel layer.

[0058] The application layer can include a series of application packages. Figure 3 As shown, the application package can include applications such as camera and gallery.

[0059] In some embodiments, the camera is used to capture videos or images in response to user operations, and the gallery is used to store videos or images captured by the camera.

[0060] The application framework layer provides application programming interface (API) and programming framework for the applications in the application layer. The application framework layer includes some predefined functions. Figure 3As shown, some shooting-related modules in the application framework layer may include a camera framework (CameraFwk), a media recorder (Media Recorder), and a media provider (Media Provider).

[0061] The Android Runtime includes core libraries and a virtual machine. The Android Runtime is responsible for scheduling and management of the Android system. In some embodiments of the present application, an application cold start will be executed within the Android Runtime. The Android Runtime can then obtain the application's optimization file status parameters. The Android Runtime can then use the optimization file status parameters to determine whether the optimization file has become outdated due to a system upgrade and return the determination result to the application management module.

[0062] The system library can include multiple functional modules, such as a surface manager, a 3D graphics processing library (e.g., OpenGL ES), a 2D graphics engine (e.g., SGL), and a video recording service.

[0063] The surface manager manages the display subsystem and provides fusion of 2D and 3D layers for multiple applications. The 3D graphics library implements 3D graphics drawing, image rendering, compositing, and layer processing. The 2D graphics engine is a drawing engine for 2D drawing. The recording service receives video streams captured by the recording function and stores them in the electronic device's memory.

[0064] The HAL includes the camera HAL and algorithm modules related to the embodiments of the present application.

[0065] The kernel layer is the layer between hardware and software. The kernel layer includes at least the display driver, camera driver, audio driver, sensor driver, etc. The camera HAL controls the camera by calling the camera driver.

[0066] It should be noted that although the embodiments of the present application are described using the Android system as an example, its basic principles are also applicable to electronic devices based on operating systems such as iOS and Windows.

[0067] based on Figure 3 The software framework shown in the figure shows the implementation process of the electronic device taking images as follows:

[0068] In response to a user's image capture operation, the camera invokes the camera through modules such as the camera framework, camera HAL, and camera driver. The camera captures multiple frames of images in response to the camera call. The captured multiple frames of images may be processed by an algorithm module and transmitted to the camera after processing. In an embodiment of the present application, the camera generates a capture result based on the multiple frames of images in response to the user's capture operation, and the capture result is stored in a gallery.

[0069] That is, the camera calls the camera in response to a user's single capture operation, and the camera captures multiple frames of images in response to the call. Based on the multiple frames of images, the electronic device generates an image that is stored in the image gallery in response to the user's capture. Therefore, the multiple frames of images are not the capture results stored in the image gallery by the electronic device in response to the user's capture.

[0070] In the process of the camera generating the shooting results, the image processing method provided in the following embodiments of the present application is used to achieve the purpose of removing stripes in the image.

[0071] Figure 4 An image processing method provided in the embodiment of the present application is applied in the following scenarios: Figure 2 Taking the camera shown as an example, the camera responds to the camera call and sequentially acquires multiple frames of images with the same exposure time (first exposure time). The object being photographed is in an environment with a stroboscopic light source and is stationary. The camera's exposure time (i.e., the camera's exposure time) is less than the stroboscopic period of the stroboscopic light source (i.e., the stroboscopic period). The stroboscopic period is the duration between adjacent bright (i.e., lit) and dark (i.e., extinguished) periods of the light source. For example, a fluorescent lamp uses 220V and 50Hz AC power, and the stroboscopic period is 10 milliseconds. The camera's exposure time is less than 10 milliseconds.

[0072] Based on the aforementioned scenario definition, it can be understood that the multiple frames captured by the camera have different timestamps and the same exposure, with the timestamp indicating the time the image was captured. Because the subject is stationary (i.e., not moving), there is no displacement between the images of the subject in each frame. Furthermore, because the exposure duration is shorter than the flash duration of the stroboscopic light source, the multiple frames exhibit banding.

[0073] Figure 4 In the diagram, the dashed line with an arrow represents the data flow, and the solid line with an arrow represents the step flow. Figure 4 The following steps are included:

[0074] S11. Perform a mean operation on multiple frame images to obtain a multi-frame fused image Frame_GT (which may be referred to as a fused image for short).

[0075] The mean operation refers to adding multiple frames of images and then averaging the added results to obtain a fused image.

[0076] Based on the above scenario, because the camera's exposure time is shorter than the flashing period of the stroboscopic light source, each frame of image captured by the camera in the stroboscopic light environment includes stripes. It can be understood that each frame of image may have local stripes or global stripes. Local stripes refer to the presence of stripes in part of the image area, and global stripes refer to the presence of stripes in the entire image area.

[0077] Since the multiple frames are taken at different times, the positions of the stripes in the multiple frames are different. Therefore, a fused image without stripes can be obtained by performing a mean operation on the multiple frames.

[0078] In some implementations, the multiple frames include a first frame (denoted as Frame_0), a second frame (denoted as Frame_1), ..., an N-1th frame (denoted as Frame_N-1), and an Nth frame Frame_N. Frame_0, Frame_1, ..., Frame_N-1 are added together and averaged to obtain a fused image Frame_GT.

[0079] S12. Obtain the difference between the fused image and the image to be processed to obtain a banding mask image banding_mask.

[0080] It is understandable that the image to be processed also has stripes.

[0081] In some implementations, the image to be processed is Frame_N, that is, the image to be processed is a newly captured frame, and the fused image is acquired based on an image frame captured before the newly captured frame.

[0082] In some implementations, the difference between Frame_GT and the image to be processed Frame_N is calculated to obtain a strip mask image.

[0083] Because there are no stripes in the fused image, the difference image is a striping mask image, recorded as banding_mask. The striping mask image can be understood as an image that contains bright strips and dark strips but no other content.

[0084] S13 . Extracting a dark stripe mask image and a bright stripe mask image from the stripe mask image banding_mask to obtain a stripe mapping image BandingMap.

[0085] In some implementations, a threshold value thr_1 is set, and pixel values ​​in banding_mask that are greater than or equal to the threshold value thr are retained unchanged (or set to 1), and pixel values ​​less than thr_1 are all set to 0, thereby obtaining a dark stripe mask image banding_dark_mask. The inverted image of banding_mask is obtained, for example, after multiplying banding_mask by -1, the pixel values ​​greater than or equal to the threshold value thr_1 are retained unchanged (or set to 1), and the pixel values ​​less than thr_1 are all set to 0, thereby obtaining a light stripe mask image banding_light_mask. It can be understood that the dark stripe mask image banding_dark_mask only contains dark stripes, and the light stripe mask image banding_light_mask only contains light stripes. The dark stripe mask image banding_dark_mask and the light stripe mask image banding_light_mask are collectively referred to as BandingMap.

[0086] Based on the above acquisition method, the BandingMap may be a grayscale image or a binary image.

[0087] S14. According to BandingMap, the fused image Frame_GT and the image to be processed are fused to obtain the image T without stripes.

[0088] The principle of fusion is that for each pixel in BandingMap, the larger the pixel value, the greater the difference between Frame_GT and the image to be processed at the pixel, which means that the pixel is likely to be a pixel on the strip. Therefore, the value of the pixel is mainly obtained based on the value of Frame_GT at the pixel. Conversely, the smaller the pixel value, the less likely the pixel is to be a pixel on the strip. The value of the pixel is mainly obtained based on the value of the image to be processed at the pixel, thereby achieving the purpose of removing the stripes.

[0089] Specifically, a pixel in BandingMap is called P, with coordinates (x1, y1). The pixel with coordinates (x1, y1) in Frame_GT is called P1, the pixel with coordinates (x1, y1) in the image to be processed Frame_N is called P2, and the pixel with coordinates (x1, y1) in the stripped image T is called P3. P3 is obtained based on the weighted fusion (e.g., weighted sum) of P1 and P2. The larger the value of P, the greater the weight of P1 and the smaller the weight of P2.

[0090] That is to say, when expanded to each pixel point, the larger the value of a pixel point in BandingMap, the greater the weight of Frame_GT at the pixel point, and the smaller the weight of Frame_N at the pixel point.

[0091] The specific implementation of the fusion will be described in detail in the following embodiments.

[0092] In this embodiment, unlike the traditional solution of adjusting the camera exposure time to prevent the occurrence of stripes, under the premise that stripes have already appeared, BandingMap is used to determine the contribution of the strip-free image and the image to be processed to the pixels on the image. That is, the strip-free pixels are used to correct or replace the striped pixels, thereby achieving the purpose of removing the stripes.

[0093] Figure 4 The process shown is performed under the assumption that the object being photographed is motionless. In practice, the object being photographed may be in motion. In order to better remove the blur caused by motion (referred to as motion blur), the camera configured in the electronic device includes a stagger HDR camera.

[0094] The stagger HDR camera can simultaneously capture multiple frames of images with different exposure times in a single shot.

[0095] In the following embodiments of the present application, the image frames obtained using a longer exposure time are referred to as long frames (abbreviated as N frames), and the image frames obtained using a shorter exposure time are referred to as short frames (abbreviated as S frames). The longer exposure time is greater than the stroboscopic period of the light source, and the shorter exposure time is less than the stroboscopic period of the light source. For example, a fluorescent lamp uses 220V and 50 Hz AC power, and the stroboscopic period is 10 milliseconds. The stagger HDR camera is configured with a longer exposure time greater than or equal to 10 milliseconds, and a shorter exposure time less than 10 milliseconds. Therefore, long frames taken under fluorescent lights will not have stripes, but short frames will have stripes. Figure 1 For example.

[0096] It is understandable that when the subject is moving, the long frame obtained by using a longer exposure time will have motion blur, while the short frame obtained by using a shorter exposure time will not have motion blur.

[0097] For any frame, the time it was captured is used as its acquisition timestamp. N frames acquired in descending order of acquisition timestamps are denoted as N1, N2, and so on. N and S frames with the same acquisition timestamp (i.e., the same capture time) are called time-aligned long and short frames.

[0098] Figure 5Another image processing method is provided in an embodiment of the present application. In this embodiment, the input is a long frame and a short frame. The long frame and the short frame as input are two time-aligned frames, N1 and S1 being used as an example here. The output is an image T without stripes.

[0099] Figure 5 In the diagram, the dashed line with an arrow represents the data flow, and the solid line with an arrow represents the step flow. Figure 5 The following steps are included:

[0100] S21 , registering the time-aligned long frame and short frame to obtain a registered long frame and a registered short frame.

[0101] The purpose of registration is to align pixel positions. Aligning pixel positions can be understood as aligning pixels representing the same physical location. This means obtaining a correspondence between pixels in the long and short frames. This correspondence is based on the actual position of the subject. For example, a point in real space is represented by pixel a in the long frame and pixel b in the short frame. Registration involves establishing a correspondence between pixel a and pixel b. Registration can be performed in a variety of ways, which will not be detailed here.

[0102] S22 , performing brightness alignment on the registered long frame and the registered short frame to obtain a brightness-processed long frame.

[0103] In some implementations, the short frame is used as a reference frame, and a wrap calculation is performed on the long frame to obtain a brightness-processed long frame.

[0104] S23. Calculate the difference BandingMap between the long frame and the short frame.

[0105] It can be understood that, since the long frame has no stripes but the short frame has stripes, the BandingMap at least characterizes the difference caused by the stripes between the brightness-processed long frame and the registered short frame.

[0106] In some implementations, in order to more accurately reflect the differences caused by banding, the difference between the brightness-processed long frame and the registered short frame is calculated as the BandingMap.

[0107] BandingMap at least characterizes the difference caused by banding between the brightness-processed long frame and the registered short frame, while removing the difference caused by the lack of position alignment and brightness alignment.

[0108] In the case of subject motion, the BandingMap also represents the difference caused by motion between the brightness-processed long frame and the registered short frame.

[0109] BandingMap can be understood as an image, where the larger the value of a pixel is, the greater the difference between the brightness processed long frame and the short frame after registration at that pixel.

[0110] The specific implementation of S23 will be described in detail in the following embodiments.

[0111] S24. According to the BandingMap, the brightness-processed long frame and the registered short frame are fused to obtain an image T without stripes.

[0112] The principle of fusion is that for each pixel in the BandingMap, the larger the pixel value, the greater the difference between the long frame and the short frame after registration and brightness alignment at the pixel, which means that the pixel is likely to be a pixel on the strip. Therefore, the value of the pixel is mainly obtained based on the value of the long frame at the pixel. Conversely, the smaller the pixel value, the less likely the pixel is to be a pixel on the strip, and the value of the pixel is mainly obtained based on the value of the short frame at the pixel, thereby achieving the purpose of removing the stripes.

[0113] The specific implementation of the fusion will be described in detail in the following embodiments.

[0114] In this embodiment, unlike the traditional solution of adjusting the camera exposure time to prevent the occurrence of stripes, when stripes have already appeared, the information of the long frame is used to compensate for the stripes that appear on the short frame, thereby achieving the purpose of removing the stripes.

[0115] Figure 6 This is the specific process of S23 obtaining BandingMap. Figure 6 The following steps are included:

[0116] S231. Calculate the absolute difference between adjacent long frames to obtain a motion alpha map.

[0117] It is understandable that adjacent long frames have different timestamps and the same exposure duration. When the subject is moving, the absolute difference between adjacent long frames represents the difference caused by the motion of the subject, so the motion alpha map represents the difference caused by the motion of the subject.

[0118] In some implementations, in order to save computing resources and obtain more accurate image processing results, Figure 5 The input is N1 frame and S1 frame, so the adjacent long frames here are N1 frame and N2 frame.

[0119] In this embodiment, it is assumed that the N2 frame is compared with the N1 frame, and the object being photographed has moved. Figure 6 As shown, the head of the person being photographed is twisted in frame N2 relative to frame N1.

[0120] S232: Binarize the motion Alpha map to obtain a motion mask motion_mask.

[0121] In some implementations, a threshold value thr_1 is set, and pixels above or equal to the threshold value thr_1 are considered to be pixels in the motion area, and the pixel value is set to 0. Pixels below the threshold value thr_1 are considered to be pixels in the static area, and the pixel value is set to 1. Therefore, in motion_mask, 0 indicates a pixel with motion blur, and 1 indicates a pixel without motion blur.

[0122] S233. Calculate the difference between the long frame and the short frame with the same timestamp to obtain diff_mask.

[0123] The difference between long and short frames with the same timestamp represents a stripe.

[0124] Combine Figure 5 As shown, the long frame with the same timestamp is N1 and the short frame is S1.

[0125] S234. Multiply motion_mask and diff_mask to obtain banding_mask.

[0126] Based on the rules of pixel value settings in motion_mask, the pixels with a pixel value of 0 in banding_mask are pixels where motion blur is set to 0, and the pixels with a pixel value of 1 are pixels where there is no motion blur. Therefore, the pixels affected by motion are eliminated from banding_mask, and only the pixels on the light and dark strips are included.

[0127] S235 . Extract the dark band banding_dark_mask and the bright band banding_light_mask from banding_mask to obtain BandingMap.

[0128] Binarize the banding_mask to get the BandingMap.

[0129] In some implementations, a threshold value thr_2 is set, and pixels in the banding_mask greater than or equal to the threshold value thr_2 are retained at their original values, while pixels less than thr_2 are set to 0, resulting in a banding_dark_mask. The banding_mask is then multiplied by -1, and pixels greater than or equal to the threshold value thr_2 are retained at their original values, while pixels less than thr_2 are set to 0, resulting in a banding_light_mask. The banding_dark_mask and banding_light_mask are collectively referred to as a BandingMap. It will be appreciated that this approach yields a grayscale BandingMap.

[0130] In other implementations, pixels in banding_mask greater than or equal to a threshold value thr_2 are set to 1, and pixels less than thr_2 are set to 0 to obtain banding_dark_mask. Banding_mask is then multiplied by -1, and pixels greater than or equal to the threshold value thr_2 are set to 1, and pixels less than thr_2 are set to 0 to obtain banding_light_mask. This method obtains a binary BandingMap.

[0131] It is understandable that, whether it is a grayscale image or a binary image, in the BandingMap, 0 indicates that the pixel is not on the stripe, and a non-zero value indicates that the pixel is on the stripe.

[0132] Figure 6 The BandingMap acquisition method shown in the figure is based on the characteristics of long and short frames captured by the stagger HDR camera. It can not only obtain stripes, but also obtain stripes without the influence of motion, laying the foundation for accurately removing stripes from the image to be processed.

[0133] It is understandable that Figure 6 The inputs of the difference calculation step S1 and N1 are only examples. An alternative method is to replace N1 with the aforementioned brightness-processed long frame obtained by N1 transformation, and replace S1 with the aforementioned registered short frame obtained by S1 transformation.

[0134] from Figure 4 and Figure 5 It can be seen that in the image processing method provided by the embodiment of the present application, in the fusion step (S14 and S24), the fused image is the image to be processed, that is, the image to be striped and the image without stripes. Figure 4 Frame_N in or Figure 5 After registration, the short frame in the image without stripes is as follows Figure 4Frame_GT or Figure 5 Luminance processing in long frames.

[0135] The basis of fusion is BandingMap. BandingMap determines whether the pixels in the fused result image are mainly based on the image to be processed or the image without banding.

[0136] The specific method of fusion will be explained in detail below.

[0137] Figure 7 This is a specific process of a fusion step in the image processing method provided in the embodiment of the present application. Figure 5 Based on the process shown, S24 includes the following steps:

[0138] S241 : Based on the BandingMap and a preset first rule, obtain a first type of weight and a second type of weight.

[0139] As mentioned above, in the BandingMap, 0 indicates that the pixel is not on the strip, and a non-zero value indicates that the pixel is on the strip. The first rule stipulates that the larger the value in the BandingMap, the greater the first-category weight and the smaller the second-category weight. The first-category weight is the weight of the long frame after brightness processing, and the second-category weight is the weight of the short frame after registration. The sum of the first and second-category weights is 1.

[0140] For a binary BandingMap, 1 corresponds to the pre-configured first and second weights, and 0 corresponds to the pre-configured third and fourth weights. The first and third weights are first-class weights, and the second and fourth weights are second-class weights. Assuming the value of a pixel P in the BandingMap is 1, and the position coordinates of pixel P in the BandingMap are (x, y), then pixel P corresponds to the first and second weights. The first weight represents the weight of the pixel with coordinates (x, y) in the long frame of brightness processing, and the second weight represents the weight of the pixel with coordinates (x, y) in the short frame after registration.

[0141] Similarly, obtain the first-category weight and second-category weight corresponding to each pixel in the BandingMap.

[0142] For grayscale BandingMap, the corresponding relationship between non-zero and the first type of weight is as follows Figure 8 As shown, Figure 8 In , the horizontal axis represents the size of BandingMap (i.e. Banding strength), and the vertical axis represents the first-class weight corresponding to BandingMap. Figure 8 It can be seen that within a certain range [X min , X max], the larger the value, the larger the corresponding first-class weight. The first-class weight value can be based on the preset [Y min , Y max ] is obtained by interpolation calculation and other methods. For any pixel, the sum of the first type of weight and the second type of weight is 1, so after obtaining the first type of weight, the second type of weight can be obtained.

[0143] S242 : Based on the first type of weights and the second type of weights, a weighted sum of the brightness-processed long frame and the registered short frame is calculated to obtain an image T without stripes.

[0144] Figure 9 For Figure 5 Another specific process of fusion calculation proposed based on Figure 7 The difference between the illustrated process and the previous one is that a weight mapping network is used instead of the first rule to obtain the first and second weights. It is understood that the weight mapping network is pre-trained, and the specific training process is not detailed here. After training, the weight mapping network learns the following: the larger the value in the BandingMap, the larger the first weight, and the smaller the second weight. The first weight is the weight for the long frame processed by luminance, and the second weight is the weight for the short frame after registration.

[0145] Figure 10 For Figure 5 Another specific process of fusion calculation proposed based on Figure 7 and Figure 9 The difference is that instead of obtaining weights separately, the BandingMap and the image to be fused are fed into the fusion network together, generating the fusion network output. The image to be fused consists of a long frame processed for brightness and a short frame after registration. The fusion network is pre-trained and has the following capabilities: it uses the BandingMap as prior information to guide the fusion of the images to be fused. Pixels with larger values ​​in the BandingMap have larger first-category weights and smaller second-category weights. The first-category weights are for the long frame processed for brightness, while the second-category weights are for the short frame after registration.

[0146] The above fusion methods are all based on BandingMap. The larger the value in BandingMap, the greater the difference between the corresponding pixel in the long frame and the short frame. Therefore, the more likely it is a pixel on the strip, so the pixel value in the long frame is mainly used as a reference during fusion. Conversely, the more likely it is not a pixel on the strip, the pixel value in the short frame is mainly used as a reference during fusion. The fusion result can not only remove the strips, but also has the advantages of the short frame.

[0147] The above Figure 5 The fusion method is explained based on the process shown in Figure 4Based on the above, the specific method of S14 can be found in Figure 7 S241 and S242 shown, or Figure 9 S241 and S242 shown, or Figure 10 The S241 shown is not described here in detail.

[0148] It is understandable that the image processing method provided in the embodiment of the present application is based on BandingMap and utilizes long frames without stripes ( Figure 4 The average of the different image frames in the image can be considered the long frame) is used to supplement the pixels in the short frames to remove the stripes. This method processes images that have stripes to remove them, rather than preventing them. Therefore, there is no need to adjust the exposure time in advance to prevent stripes.

[0149] An embodiment of the present application further provides a computer-readable storage medium having instructions stored thereon. When the instructions are executed on an electronic device, the electronic device executes the image processing method provided in the above embodiment.

[0150] An embodiment of the present application further provides a computer program product. When the computer program product is run on an electronic device, the electronic device implements the image processing method provided in the above embodiment.

[0151] The above is only a specific embodiment of the present application, but the scope of protection of this application is not limited to this. Any changes or substitutions within the technical scope disclosed in this application should be included in the scope of protection of this application. Therefore, the scope of protection of this application should be based on the scope of protection of the claims.

Claims

1. An image processing method, characterized in that: Applied to an electronic device, a camera running in the electronic device acquires multiple frames of images by calling a camera to shoot a subject with a first exposure duration, the subject being in an environment with a stroboscopic light source and being stationary, and the first exposure duration being less than a stroboscopic period. The method includes: Performing a mean operation on the multiple frames of images to obtain a fused image; Based on the difference between the fused image and the image to be processed, a strip mapping image is obtained; the multiple frames of images include the image to be processed; Based on the strip mapping image, the fused image and the image to be processed are fused to obtain an image without strips. During the fusion process, the value of the first pixel in the image without strips is obtained based on a weighted fusion method of the value of the first pixel in the fused image and the value of the first pixel in the image to be processed. The larger the value of the first pixel in the strip mapping image, the greater the weight of the fused image at the first pixel, and the smaller the weight of the image to be processed at the first pixel. The first pixel is any pixel.

2. The method according to claim 1, characterized in that The step of obtaining a strip mapping image based on a difference between the fused image and the image to be processed includes: Obtaining a difference between the fused image and the image to be processed to obtain a strip mask image; A dark stripe mask image and a bright stripe mask image are extracted from the stripe mask image to obtain the stripe mapping image.

3. The method according to claim 2, characterized in that Extracting a dark stripe mask image and a bright stripe mask image from the stripe mask image includes: retaining the original values ​​or setting the pixel values ​​in the stripe mask image that are greater than or equal to a first threshold to 1, and setting the pixel values ​​that are less than the first threshold to 0, to obtain the dark stripe mask image; The pixel values ​​in the flipped image of the stripe mask image that are greater than or equal to the first threshold are retained as original values ​​or set to 1, and the pixel values ​​that are less than the first threshold are set to 0, to obtain the bright stripe mask image.

4. An image processing method, characterized in that: The method is applied to an electronic device, wherein a camera running in the electronic device captures an object by calling a stagger HDR camera, wherein the object is in an environment with a stroboscopic light source, and the exposure time of a long frame captured by the stagger HDR camera is greater than or equal to the stroboscopic period, and the exposure time of a short frame captured by the stagger HDR camera is less than the stroboscopic period. The method comprises: acquiring a stripe mapping image based on a difference between a first long frame and a first short frame, wherein the first long frame and the first short frame have the same shooting time; Based on the strip mapping image, a first image and a second image are fused to obtain a strip-removed image, wherein the first image is obtained based on the first long frame, and the second image is obtained based on the first short frame. During the fusion process, a value of a first pixel in the strip-removed image is obtained based on a weighted fusion method of a value of the first pixel in the first image and a value of the first pixel in the second image. The larger the value of the first pixel in the strip mapping image, the greater the weight of the first image in the first pixel, and the smaller the weight of the second image in the first pixel.

5. The method according to claim 4, characterized in that The step of acquiring a stripe mapping image based on a difference between the first long frame and the first short frame includes: calculating a difference between the first long frame and the first short frame to obtain a difference mask image, where the difference mask image represents a difference between the first long frame and the first short frame caused by motion and striping of the object; acquiring a stripe mask image based on a pre-acquired motion mask image and the difference mask image, wherein the motion mask image represents a difference between the first long frame and an adjacent long frame caused by the motion of the object, and the stripe mask image represents a difference between the first long frame and the first short frame caused by stripes; A dark stripe mask image and a bright stripe mask image are extracted from the stripe mask image to obtain the stripe mapping image.

6. The method according to claim 5, characterized in that Before acquiring the stripe mask image based on the pre-acquired motion mask image and the difference mask image, the method further includes: A motion mask image is acquired based on a difference between the first long frame and the second long frame, wherein the motion mask image represents a difference between the first long frame and the second long frame caused by the motion of the object, and the first long frame and the second long frame are image frames with adjacent time stamps.

7. The method according to claim 5 or 6, characterized in that Extracting a dark stripe mask image and a bright stripe mask image from the stripe mask image includes: retaining the original values ​​or setting the pixel values ​​in the stripe mask image that are greater than or equal to a first threshold to 1, and setting the pixel values ​​that are less than the first threshold to 0, to obtain the dark stripe mask image; The pixel values ​​in the inverted image of the stripe mask image that are greater than or equal to the first threshold are retained as original values ​​or set to 1, and the pixel values ​​that are less than the first threshold are set to 0, to obtain the bright stripe mask image.

8. The method according to any one of claims 4 to 7, characterized in that: Before fusing the first image and the second image based on the strip mapping image, the method further includes: Registering the first long frame and the first short frame to obtain a registered long frame and a registered short frame; Brightness alignment is performed on the registered long frame to the registered short frame to obtain a brightness-processed long frame, the first image is the brightness-processed long frame, and the second image is the registered short frame.

9. An electronic device, characterized in that: The electronic device includes: a camera, one or more processors, a memory and a touch screen; The camera captures a subject with a first exposure duration to obtain multiple frames of images, where the first exposure duration is less than a stroboscopic period of a stroboscopic light source; or the camera is a staggered high dynamic range (HDR) camera, where the exposure duration of long frames captured by the stagger HDR camera is greater than or equal to the stroboscopic period, and the exposure duration of short frames captured by the stagger HDR camera is less than the stroboscopic period; The memory is used to store program code; the processor is used to run the program code, so that the electronic device implements the image processing method according to any one of claims 1 to 8.

10. A computer-readable storage medium, characterized in that Instructions are stored thereon, and when the instructions are executed on an electronic device, the electronic device executes the image processing method according to any one of claims 1 to 8.

11. A computer program product, characterized in that When the computer program product is run on an electronic device, the electronic device is enabled to implement the image processing method according to any one of claims 1 to 8.

Citation Information

Patent Citations

  • Image processing method, device and equipment

    CN109671106A

  • WDR imaging with LED flicker mitigation

    CN109845242A

  • Stroboscopic detection method and related device

    CN114630106A

  • Image processing method, system and camera

    CN115314627A

  • HDR image sensor with LFM and reduced motion blur

    US20200092459A1