Image blur processing method and electronic device

By combining binocular cameras and ToF cameras to obtain high-precision depth information in electronic devices for dense processing and blurring processing, the problem of poor image blurring effect in the prior art is solved, and more accurate front and back scene segmentation and natural blurring effect are achieved.

CN112614057BActive Publication Date: 2025-08-15HUAWEI TECH CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN201910880232.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2019-09-18
Publication Date
2025-08-15
Estimated Expiration
2039-09-18

AI Technical Summary

Technical Problem

In the prior art, electronic devices have poor effects in image blurring processing, and there are problems such as false blurring of the foreground, poor edge effects of the foreground, or flickering backgrounds.

Method used

The method of combining binocular cameras and time-of-flight ToF cameras is adopted to densely process the RGB images by obtaining high-precision depth information, improving the accuracy of front and back scene segmentation, and blurring according to the depth value, and optimizing edge recognition with convolutional neural network.

Benefits of technology

It improves the image blur effect, especially in scenes with similar textures in front and back scenes, which can achieve good edge segmentation effect, reduce background flickering, and enhance visual experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN112614057B_ABST
    Figure CN112614057B_ABST
Patent Text Reader

Abstract

The present application provides an image blur processing method and electronic device for optimizing the image blur effect. The method is applied to an electronic device, wherein the electronic device is provided with a first camera and a second camera, wherein the first camera is a binocular camera or a multi-camera camera, and the second camera is a ToF camera or a structured light camera. The method comprises: simultaneously starting the first camera and the second camera to perform a shooting operation on a first scene, wherein the first camera obtains an RGB image and a first depth image corresponding to the first scene, and the second camera obtains depth information of the first scene; using the depth information of the first scene to perform densification processing on the first depth image to obtain a second depth image, wherein the number of invalid pixels in the second depth image is less than the number of invalid pixels in the first depth image; using the second depth image to perform foreground and background segmentation on the RGB image; and performing blur processing on the background image or the foreground image in the RGB image to obtain an RGB image with a background or foreground blur effect.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of terminal technology, and in particular to an image blurring processing method and an electronic device. Background Art

[0002] Taking images with a blurred background effect can highlight the key objects in the image and ignore other supporting objects. It is often used in portraits, artworks and other photography.

[0003] Currently, widely used electronic devices with camera functions, such as mobile phones and tablets, can achieve a certain blurring effect by simulating the large aperture processing of SLR cameras on the captured images in the later stage of image shooting. However, the existing technology has a poor blurring effect on the image, and there are problems such as the foreground being mistakenly blurred, poor foreground edge effect or background flickering. Summary of the Invention

[0004] The embodiments of the present application provide an image blur processing method and an electronic device, which are used to solve the technical problem that electronic devices in the prior art have poor blur processing effects on images.

[0005] In a first aspect, an embodiment of the present application provides an image blur processing method, which is applied to an electronic device, wherein a first camera and a second camera are provided on the electronic device, wherein the first camera is a binocular camera or a multi-camera camera, and the second camera is a time-of-flight ToF camera or a structured light camera. The method includes: in response to a first instruction, starting the first camera and the second camera to perform a shooting operation on a first scene, the first camera obtains a red, green, and blue (RGB) image and a first depth image corresponding to the first scene, and the second camera obtains depth information of the first scene; using the depth information of the first scene to perform densification processing on the first depth image to supplement invalid pixels in the first depth image to obtain a second depth image, wherein the number of invalid pixels in the second depth image is less than the number of invalid pixels in the first depth image; using the second depth image to perform foreground and background segmentation on the RGB image to obtain a foreground image and a background image; blurring the background image or the foreground image to obtain an RGB image with a background or foreground blur effect.

[0006] In an embodiment of the present application, when a first camera captures a first depth image of a first scene, a second camera is simultaneously controlled to capture the first scene. An RGB image and a first depth image of the first scene are obtained based on the first camera, and high-precision depth information of the first scene is obtained based on the second camera. The high-precision depth information obtained by the second camera is then used to perform densification processing on the first depth image captured by the first camera to obtain a second depth image. The precision of the second depth image is higher than that of the first depth image (that is, the number of invalid pixels in the second depth image is less than that in the first depth image). The high-precision second depth image is then used to segment the foreground and background of the RGB image. This can improve the accuracy of edge recognition of the foreground and background of the RGB image. Even in scenes where the texture, color, etc. of the foreground and background of the image are similar, a good edge segmentation effect can be obtained, thereby improving the blurring effect of the image.

[0007] In one possible design, when segmenting the foreground and background of RGB, the foreground and background in the RGB image can also be segmented based on the second depth image in combination with a trained convolutional neural network (CNN) model; wherein the input of the trained CNN model is an RGB image in which the subject is a specific type of subject, such as a portrait, and the output of the trained CNN model is the image portion corresponding to the subject in the image, i.e., a portrait image.

[0008] In this embodiment, when the subject being photographed is an image of a specific type of subject, CNN is further combined to identify the specific type of subject being photographed, intelligently segment the foreground and background, further improve the accuracy of foreground and background segmentation, and optimize the foreground and background edge segmentation effect to achieve different blurring effects.

[0009] In one possible design, the background image can be blurred according to the second depth image, so that the pixel area with a larger depth value in the background image corresponds to a higher degree of blurring, and then the foreground image and the blurred background image are fused to obtain an RGB image with a background blur effect.

[0010] In this way, the greater the depth of field in the RGB image background, the higher the blur intensity can be, making the blur effect closer to the real optical defocus effect, thereby improving the user's visual experience.

[0011] In another possible design, the foreground image can also be blurred according to the second depth image, so that the pixel area with a smaller depth value in the foreground image corresponds to a higher degree of blurring, and then the background image and the blurred foreground image are fused to obtain an RGB image with a foreground blur effect.

[0012] In this way, the intensity of blurring can be made higher in areas with smaller depth of field in the foreground of the RGB image, thereby achieving different blurring effects and improving the user's visual experience.

[0013] In one possible design, after obtaining the RGB image with the background or foreground blur effect, the RGB image with the background or foreground blur effect can be weighted fused with the original RGB image. This can effectively eliminate image blur and moiré issues.

[0014] In one possible design, the technical solution of the embodiment of the present application can be used in a scene where a video image is blurred. Specifically, the first camera can capture a video corresponding to the first scene, the video including multiple consecutive RGB images, and obtain a first depth image corresponding to each RGB image in the multiple consecutive RGB images. The second camera can obtain depth information corresponding to each RGB image in the multiple consecutive RGB images. Then, for each RGB image in the multiple consecutive RGB images, the first depth image corresponding to the RGB image is densified using the depth information corresponding to the RGB image obtained by the second camera to supplement invalid pixels in the first depth image corresponding to the RGB image, thereby obtaining a second depth image corresponding to the RGB image. Then, the second depth image corresponding to each RGB image in the multiple consecutive RGB images can be used to perform foreground and background segmentation on the RGB image to obtain a foreground image and a background image corresponding to the RGB image. Finally, the foreground image or background image corresponding to each RGB image in the multiple consecutive RGB images can be blurred to obtain a depth image with a blurred background or foreground effect corresponding to each RGB image.

[0015] The above solution, based on the video images captured by the first camera, combines the second camera to obtain depth information of the captured scene. The high-precision depth information obtained by the second camera is used to densify the first depth image corresponding to each frame of the RGB image captured by the first camera and segment the foreground and background. This improves the accuracy of video foreground edge recognition. Even in scenes with similar textures and colors of the foreground and background, good edge segmentation effects can be obtained, which can improve the visual experience of video blur.

[0016] In one possible design, before blurring the foreground image or background image corresponding to each RGB image in the multiple continuous RGB images, N adjacent RGB images in the multiple continuous RGB images may be weightedly fused, where N is a positive integer greater than 1, and the weighting coefficient corresponding to each RGB image in the N adjacent RGB images is determined based on the depth information corresponding to the RGB image frame acquired by the second camera and / or the foreground and background edge information of the RGB image frame identified by the CNN. Therefore, when blurring the background image or the foreground image, the background image or foreground image corresponding to the weighted fused RGB image may be specifically blurred.

[0017] In this embodiment, since the depth information between adjacent video frames obtained by the second camera has good stability and the edge information between adjacent RGB images detected by CNN has good stability, adjacent video frames are weightedly fused based on the depth information corresponding to the adjacent RGB images obtained by the second camera and / or the foreground and background edge information of the RGB images identified by CNN. This can make the image transition between the adjacent RGB images after the weighted fusion process more natural, thereby improving the problem of background flickering in the video frames captured by the binocular camera in the prior art when shooting videos with a blur effect due to the poor inter-frame stability of the binocular camera.

[0018] According to a second aspect, an electronic device is provided, comprising a first camera, a second camera, and at least one processor, wherein the first camera is a binocular camera or a multi-camera camera, and the second camera is a time-of-flight (ToF) camera or a structured light camera; the processor is configured to, in response to a first instruction, start the first camera and the second camera to perform a shooting operation on a first scene; the first camera is configured to perform a shooting operation on the first scene to obtain a red, green, and blue (RGB) image and a first depth image corresponding to the first scene; the second camera is configured to perform a shooting operation on the first scene to obtain depth information of the first scene; the processor is further configured to use the depth information of the first scene to perform densification processing on the first depth image to supplement invalid pixels in the first depth image to obtain a second depth image, wherein the number of invalid pixels in the second depth image is less than the number of invalid pixels in the first depth image; use the second depth image to perform foreground and background segmentation on the RGB image to obtain a foreground image and a background image; and blur the background image or the foreground image to obtain an RGB image with a background or foreground blur effect.

[0019] In one possible design, when the processor uses the second depth image to segment the foreground and background of the RGB image, it may specifically use the second depth image and a trained convolutional neural network (CNN) model to segment the foreground and background in the second depth image; wherein the input of the trained CNN model is an RGB image in which the subject is a specific type of subject, and the output is the RGB image portion corresponding to the subject in the image. The specific type of subject may be a portrait.

[0020] In one possible design, when the processor blurs the background image or the foreground image, the processor may specifically: use the second depth image to blur the background image on the RGB image, wherein the pixel area with a larger depth value in the background image corresponds to a higher degree of blurring; fuse the foreground image and the blurred background image to obtain an RGB image with a background blur effect; or: blur the foreground image, wherein the pixel area with a smaller depth value in the foreground image corresponds to a higher degree of blurring; fuse the background image and the blurred foreground image to obtain an RGB image with a foreground blur effect.

[0021] In a possible design, after obtaining the RGB image with the background or foreground blur effect, the processor may further perform weighted fusion on the RGB image with the background or foreground blur effect and the RGB image.

[0022] In one possible design, the electronic device can be used to blur video images.

[0023] For example, the first camera is configured to capture a video corresponding to the first scene, the video comprising a plurality of consecutive RGB image frames, and obtain a first depth image corresponding to each of the plurality of consecutive RGB image frames. The second camera is configured to obtain depth information corresponding to each of the plurality of consecutive RGB image frames. The processor is configured to perform densification processing on the first depth image corresponding to each of the plurality of consecutive RGB image frames using the depth information corresponding to the RGB image frame obtained by the second camera, thereby supplementing invalid pixels in the first depth image corresponding to the RGB image frame to obtain a second depth image corresponding to the RGB image frame; perform foreground and background segmentation on the RGB image frame using the second depth image corresponding to each of the plurality of consecutive RGB image frames to obtain a foreground image and a background image corresponding to the RGB image frame; and perform blur processing on the foreground image or the background image corresponding to each of the plurality of consecutive RGB image frames to obtain a depth image corresponding to each of the RGB image frames with a blurred background or foreground effect.

[0024] In one possible design, before blurring the foreground image or background image corresponding to each RGB image in the multiple continuous RGB images, the processor may also perform weighted fusion on N adjacent RGB images in the multiple continuous RGB images, where N is a positive integer greater than 1, and the weighting coefficient corresponding to each RGB image in the N adjacent RGB images is determined based on the depth information corresponding to the frame RGB image acquired by the second camera and / or the foreground and background edge information of the frame RGB image identified by the CNN. Then, when blurring the background image or the foreground image, the processor may specifically blur the background image or foreground image corresponding to the weighted fused RGB image.

[0025] In a third aspect, an electronic device is provided, comprising a first camera, a second camera, at least one processor and a memory, wherein the first camera is a binocular camera or a multi-camera, and the second camera is a time-of-flight ToF camera or a structured light camera; the memory is used to store one or more computer programs; when the one or more computer programs stored in the memory are executed by the at least one processor, the electronic device is able to implement the technical solution in the first aspect of the embodiment of the present application or any possible design of the first aspect.

[0026] In a fourth aspect, a computer-readable storage medium is provided, which includes a computer program. When the computer program runs on an electronic device, the electronic device executes the technical solution in the first aspect of the embodiment of the present application or any possible design of the first aspect.

[0027] In a fifth aspect, a program product is provided, comprising instructions, which, when executed on a computer, enable the computer to execute the technical solution in the first aspect of the embodiment of the present application or any possible design of the first aspect.

[0028] In a sixth aspect, a circuit system is provided, wherein the circuit system is used to generate a first control signal, wherein the first control signal is used to control a first camera to perform a shooting operation on a first scene and obtain a first depth image corresponding to the first scene; the circuit system is also used to generate a second control signal, wherein the second control signal is used to control the second camera to perform a shooting operation on the first scene and obtain depth information of the first scene. BRIEF DESCRIPTION OF THE DRAWINGS

[0029] Figure 1 This is a schematic structural diagram of an electronic device according to an embodiment of the present application;

[0030] Figure 2 A flowchart of an image blurring processing method according to an embodiment of the present application is shown;

[0031] Figure 3 A schematic diagram of an image capture scene in an embodiment of the present application;

[0032] Figure 4A 、 Figure 4B Schematic diagrams of partial pixels of the first depth image before and after densification respectively;

[0033] Figure 5A 、 Figure 5B Schematic diagrams of the segmentation edges of RGB image and foreground and background of RGB image respectively;

[0034] Figure 6 A schematic diagram of a template corresponding to a pixel point;

[0035] Figure 7 A schematic diagram of the effect of blurring the background image;

[0036] Figure 8 Schematic diagram of another image blurring processing method in an embodiment of the present application;

[0037] Figure 9 A schematic diagram of the effect of blurring the foreground image;

[0038] Figure 10 Schematic diagram of another image blurring processing method in an embodiment of the present application;

[0039] Figure 11A This is a schematic diagram of a portrait shooting scene in an embodiment of the present application;

[0040] Figure 11B for Figure 11A A schematic diagram of an image obtained by capturing the scene shown;

[0041] Figure 11C For Figure 11B The image shown is a diagram showing the effect of blurring the background;

[0042] Figure 12 Schematic diagram of a flow chart of a method for blurring a video image. DETAILED DESCRIPTION

[0043] The technical solutions in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application.

[0044] Currently, SLR cameras capture blurred background images by directly altering the optical imaging technique using the camera's telephoto lens. This results in a clear image only within a certain object distance range, while blurring the image to varying degrees outside of this range. This effect emphasizes the key subject while ignoring other supporting elements, and is commonly used in portraits and artwork. This technique utilizes the lens's aperture, focus distance (object distance), and focal length variability to directly adjust the blur effect. A larger aperture, a closer focus distance, or a longer focal length results in a more pronounced background blur.

[0045] Currently, widely used electronic devices with camera functions, such as mobile phones and tablets, are generally equipped with small aperture lenses due to limitations in size, cost, and usage environment. Therefore, it is difficult to capture images with the same wide aperture effect as a SLR camera. In order to achieve a background blur effect for images captured by electronic devices such as mobile phones and tablets, the images can be processed by simulating a large aperture in the post-shooting process to achieve a certain blur effect. In the prior art, there are mainly two methods for post-processing the background blur of images:

[0046] The first method is to use dual cameras (binocular cameras) to simultaneously obtain two images of the same scene from different perspectives, use the principle of binocular stereo vision to calculate the depth of field information of the scene, and then use the depth of field information to separate the foreground and background of the image, and then blur the background.

[0047] However, binocular cameras are sensitive to ambient light intensity and rely heavily on the texture features of the image itself. Therefore, when the scene being photographed is far away, the scene is poorly illuminated, or lacks texture (or the texture difference between the foreground and background is small), it is difficult to accurately distinguish between the foreground and background. Therefore, the background blur images captured using this method usually have poor foreground edge effects. Furthermore, when shooting images containing multiple consecutive frames, such as animated images or videos, the stability between adjacent image frames obtained based on the binocular stereo vision principle is poor, so background flickering is a common problem.

[0048] The second method uses a single camera to capture an image, and then uses convolutional neural networks (CNN) to detect the edges of image elements of the subject (such as a portrait) in the image. The foreground and background are then separated based on the detected edges, and the background is finally blurred.

[0049] However, in the blurred image processed by this method, the degree of blurring of each pixel position in the image background is almost the same, so the sense of layering is weak and the visual effect is poor; and if the color of the subject and the background are similar, it is difficult to correctly distinguish between the foreground and the background, resulting in the problem that the subject is also blurred or the background is not blurred.

[0050] In order to solve the above-mentioned problems in the prior art, the embodiments of the present application provide an image blur processing method and an electronic device to improve the visual effect of image blur. A binocular camera and a time of flight (ToF) camera are provided on the electronic device. When the electronic device shoots an image of the first scene, the binocular camera and the ToF camera are used to shoot the first scene at the same time, and the binocular camera shoots a red-green-blue (RGB) image corresponding to the first scene, and obtains a first depth image corresponding to the first scene based on the binocular parallax of the binocular camera (wherein the value of each pixel point in the first depth image is a depth value, representing the distance between the object point corresponding to the pixel point and the electronic device). At the same time, the ToF camera is used to obtain the depth information in the first scene; then, the first depth image is processed in combination with the depth information in the first scene obtained by the ToF camera. The first depth image is densified to supplement the invalid pixels (i.e., pixels without depth values) in the first depth image to obtain a second depth image with higher precision (i.e., the second depth image has more valid pixels than the first depth image); the second depth image is then used to segment the foreground and background of the RGB image to obtain a foreground image and a background image; the background image is then blurred, wherein the greater the depth of field (i.e., the greater the depth value, i.e., the farther from the electronic device), the higher the blur intensity can be; finally, the foreground image and the virtualized background image are fused to obtain an RGB image with a background blur effect. The embodiment of the present application combines a binocular camera with a Time of Flight (ToF) camera. The ToF camera is used to obtain depth information from the first scene and then densify the first depth image captured by the binocular camera. This results in a more precise second depth image (i.e., more accurate depth information of the first scene). This second depth image is used to segment the foreground and background, thereby optimizing the accuracy of the foreground and background segmentation and, in turn, improving the edge effects of background blur. Furthermore, during the blurring process, the second depth image is used to blur the background of the RGB image, resulting in a higher blur intensity in areas with greater depth of field, making the blurring effect closer to a realistic optical defocus effect. The specific solution will be described later.

[0051] In an embodiment of the present application, if the image captured by the electronic device is a video or an animated image, the above-mentioned image blurring processing method can be performed on each frame of the video or animated image at the same time, thereby achieving an animated image or video with a background blur effect. In addition, before blurring the background of each frame of the image, the high inter-frame stability of the ToF camera (i.e., the difference between adjacent image frames obtained by the ToF camera is small and the inter-frame stability is high) can be used to perform temporal smoothing on each frame of the image to improve the problem of background flickering in the blurred images captured by the binocular camera in the prior art. The specific solution will be introduced later.

[0052] In this embodiment of the present application, if the first scene is a portrait, a CNN can be further incorporated to identify the edges of the subject and segment the background, further optimizing the foreground edge effect of the blurred image. Furthermore, when performing temporal smoothing on multiple consecutive frames, the high inter-frame stability of a CNN can be further incorporated to further mitigate background flicker and achieve a more natural transition between adjacent frames. The specific solution will be described later.

[0053] Of course, the technical concept of this application can also be applied to scenarios where the foreground of an image is blurred. Unlike shooting an image with a background blur effect, when the electronic device uses the depth information of the first scene captured by the ToF camera to separate the foreground and background in the first depth image, it blurs the foreground image. Moreover, the smaller the depth of field, the higher the blur intensity can be. Finally, the background image and the blurred foreground image are fused to obtain a depth image with a foreground blur effect.

[0054] The technical solution of the embodiment of the present application can be applied to any electronic device with an image capture function to perform an image blurring processing method. The electronic device may be, for example, a mobile phone, a mobile computer, a tablet computer, a Polaroid camera, a personal digital assistant (PDA), a media player, a smart TV, an intelligent wearable device (such as a smart watch, smart glasses and a smart bracelet), an e-reader, a handheld game console, a point of sales (POS), an in-vehicle electronic device (in-vehicle computer), etc. In an embodiment of the present application, multiple applications may be installed in the electronic device, for example, a camera application, a beauty application, a video player application, a music player application, a system setting application, a desktop application, a drawing application, a presentation application, a game application, a phone application, an email application, an instant messaging application, a photo management application, a browser application, a calendar application, a clock application, a payment application and a health management application, etc. In addition, the image blurring processing method proposed in the embodiment of the present application is applied to any image capture scene, such as taking a photo, a moving picture or a video.

[0055] The following describes a schematic structural diagram of an electronic device used in an embodiment of the present application, taking a mobile phone as an example.

[0056] like Figure 1As shown, the mobile phone 100 may include a processor 110, an external memory interface 120, an internal memory 121, a universal serial bus (USB) interface 130, a charging management module 140, a power management module 141, a battery 142, an antenna 1, an antenna 2, a mobile communication module 150, a wireless communication module 160, an audio module 170, a speaker 170A, a receiver 170B, a microphone 170C, an earphone interface 170D, a sensor module 180, a button 190, a motor 191, an indicator 192, a camera 193, a display screen 194, and a subscriber identification module (SIM) card interface 195, etc. The sensor module 180 may include a pressure sensor 180A, a gyroscope sensor 180B, an air pressure sensor 180C, a magnetic sensor 180D, an acceleration sensor 180E, a distance sensor 180F, a proximity light sensor 180G, a fingerprint sensor 180H, a temperature sensor 180J, a touch sensor 180K, an ambient light sensor 180L, a bone conduction sensor 180M, etc.

[0057] The processor 110 may include one or more processing units, for example: the processor 110 may include an application processor (AP), a modem processor, a graphics processing unit (GPU), an image signal processor (ISP), a controller, a memory, a video codec, a digital signal processor (DSP), a baseband processor, and / or a neural-network processing unit (NPU), etc. Among them, different processing units can be independent devices or integrated into one or more processors. The execution of the image blur processing method in the embodiment of the present application can be controlled by the processor 110 or completed by calling other components, such as calling the processing program of the embodiment of the present application stored in the internal memory 121, or calling the processing program of the embodiment of the present application stored in a third-party device through the external memory interface 120 to achieve post-blur processing of the captured image.

[0058] Processor 110 may also include a memory for storing instructions and data. In some embodiments, the memory in processor 110 is a cache memory. This memory can store instructions or data that have just been used or are being recycled by processor 110. If processor 110 needs to use the same instruction or data again, it can directly access the memory. This avoids duplicate accesses, reduces processor 110 latency, and thus improves system efficiency.

[0059] In some embodiments, the processor 110 may include one or more interfaces. The interfaces may include an inter-integrated circuit (I2C) interface, an inter-integrated circuit sound (I2S) interface, a pulse code modulation (PCM) interface, a universal asynchronous receiver / transmitter (UART) interface, a mobile industry processor interface (MIPI), a general-purpose input / output (GPIO) interface, a subscriber identity module (SIM) interface, and / or a universal serial bus (USB) interface.

[0060] The charging management module 140 is configured to receive charging input from a charger. The charger can be either a wireless charger or a wired charger. In some wired charging embodiments, the charging management module 140 can receive charging input from the wired charger via the USB interface 130. In some wireless charging embodiments, the charging management module 140 can receive wireless charging input via the wireless charging coil of the mobile phone 100. While charging the battery 142, the charging management module 140 can also provide power to the electronic device through the power management module 141.

[0061] The power management module 141 is used to connect the battery 142, the charging management module 140 and the processor 110. The power management module 141 receives input from the battery 142 and / or the charging management module 140, and provides power to the processor 110, the internal memory 121, the external memory, the display 194, the camera 193, and the wireless communication module 160. The power management module 141 can also be used to monitor parameters such as battery capacity, battery cycle count, and battery health status (leakage, impedance). In some other embodiments, the power management module 141 can also be set in the processor 110. In other embodiments, the power management module 141 and the charging management module 140 can also be set in the same device.

[0062] The wireless communication function of the mobile phone 100 can be implemented through the antenna 1, the antenna 2, the mobile communication module 150, the wireless communication module 160, the modem processor and the baseband processor.

[0063] Antenna 1 and Antenna 2 are used to transmit and receive electromagnetic wave signals. Each antenna in mobile phone 100 can be used to cover a single or multiple communication frequency bands. Different antennas can also be reused to improve antenna utilization. For example, antenna 1 can be reused as a diversity antenna for a wireless local area network. In other embodiments, the antennas can be used in conjunction with a tuning switch.

[0064] The mobile communication module 150 can provide solutions for wireless communications including 2G / 3G / 4G / 5G applied on the mobile phone 100. The mobile communication module 150 may include at least one filter, a switch, a power amplifier, a low noise amplifier (LNA), etc. The mobile communication module 150 can receive electromagnetic waves from the antenna 1, and filter, amplify and process the received electromagnetic waves, and transmit them to the modulation and demodulation processor for demodulation. The mobile communication module 150 can also amplify the signal modulated by the modulation and demodulation processor, and convert it into electromagnetic waves for radiation through the antenna 1. In some embodiments, at least some of the functional modules of the mobile communication module 150 can be set in the processor 110. In some embodiments, at least some of the functional modules of the mobile communication module 150 can be set in the same device as at least some of the modules of the processor 110.

[0065] The wireless communication module 160 can provide wireless communication solutions including wireless local area networks (WLAN) (such as wireless fidelity (Wi-Fi) networks), Bluetooth (BT), global navigation satellite system (GNSS), frequency modulation (FM), near field communication (NFC), infrared (IR), etc., which are applied to the mobile phone 100. The wireless communication module 160 can be one or more devices that integrate at least one communication processing module. The wireless communication module 160 receives electromagnetic waves via the antenna 2, frequency modulates and filters the electromagnetic wave signals, and sends the processed signals to the processor 110. The wireless communication module 160 can also receive the signal to be sent from the processor 110, frequency modulate it, amplify it, and convert it into electromagnetic waves for radiation through the antenna 2.

[0066] Mobile phone 100 implements display functionality through a GPU, display screen 194, and an application processor. A GPU is a microprocessor for image processing that connects display screen 194 and the application processor. The GPU is used to perform mathematical and geometric calculations for graphics rendering. Processor 110 may include one or more GPUs that execute program instructions to generate or modify display information.

[0067] The display screen 194 can be used to display information input by the user (such as text information, voice information, etc.) or information provided to the user (such as captured images, videos, etc.) and various menus of the mobile phone 100. It can also accept user input, such as user touch operations.

[0068] The display screen 194 includes a display panel. The display panel can be a liquid crystal display (LCD), an organic light-emitting diode (OLED), an active-matrix organic light-emitting diode (AMOLED), a flexible light-emitting diode (FLED), a MiniLED, a MicroLED, a Micro-oLed, a quantum dot light-emitting diode (QLED), or the like.

[0069] The display screen 194 may also include a touch panel, which is also called a touch screen, touch-sensitive screen, etc., which can collect user contact or non-contact operations on or near it (such as operations performed by the user using fingers, stylus, or any other suitable objects or accessories on or near the touch panel, which may also include somatosensory operations; the operations include single-point control operations, multi-point control operations, and other operation types), and drive the corresponding connection device according to a pre-set program.

[0070] Optionally, the touch panel may include a touch detection device and a touch controller. Wherein, the touch detection device detects the user's touch orientation and posture, and detects the signal brought by the input operation, and transmits the signal to the touch controller; the touch controller receives the touch information from the touch detection device, and converts it into information that the processor can process, and then sends it to the processor 110, and can receive the command sent by the processor 110 and execute it. In addition, the touch panel can be implemented using various types such as resistive, capacitive, infrared and surface acoustic waves, or any technology developed in the future can be used to implement the touch panel. Further, the touch panel can cover the display panel, and the user can operate on or near the touch panel covered on the display panel according to the content displayed on the display panel (the display content includes but is not limited to, soft keyboard, virtual mouse, virtual button, icon, etc.), and the touch panel detects the operation on or near it, and transmits it to the processor 110 to determine the user input, and then the processor 110 provides corresponding visual output on the display panel according to the user input.

[0071] For example, in an embodiment of the present application, after the touch detection device in the touch panel detects a touch operation input by the user, the signal corresponding to the detected touch operation is sent to the touch controller in real time; the touch controller converts the signal into touch coordinates and sends them to the processor 110; the processor 110 determines, based on the received touch coordinates, that the touch operation is specifically a click operation of clicking the "shoot" button on the shooting interface of the camera application; the processor 110 responds to the click operation input by the user, controls the binocular camera to capture the left view image and the right view image of the current scene (wherein the left view image and the right view image are both RGB images), and controls the ToF camera to obtain the depth information of the current scene; the processor 110 calculates the left view image and the right view image based on the parallax of the two cameras of the binocular camera to obtain a sparse depth image, that is, a first depth image with lower precision. ; Afterwards, the processor 110 densifies the first depth image in combination with the depth information of the first scene acquired by the ToF camera to obtain a dense depth image (having more effective pixels than a sparse depth image), that is, a second depth image with higher precision; Afterwards, the processor 110 performs foreground and background segmentation, background or foreground blurring, foreground and background fusion, etc. on the GRB image taken by the binocular camera (the GRB image can be a left view image taken by the binocular camera, or a right view image taken by the binocular camera, or an RGB image after the left view image and the right view image are fused, which is not limited in the embodiment of the present invention) to obtain an RGB image with a background or foreground blur effect; finally, the processor 110 can also control the touch panel to visually output the RGB image with a background or foreground blur effect. Optionally, when performing visual output, the second depth image can be fused to display the RGB image, such as for 3D image display.

[0072] In some embodiments, the mobile phone 100 may include 1 or N display screens 194 , where N is a positive integer greater than 1.

[0073] The mobile phone 100 can realize the shooting function through the ISP, camera 193, video codec, GPU, display 194 and application processor.

[0074] The ISP processes data fed back by camera 193. For example, when taking a photo, the shutter is opened, and light is transmitted through the lens to the camera's photosensitive element. The light signal is converted into an electrical signal, which is then passed to the ISP for processing and converted into a visible image. The ISP can also perform algorithmic optimization on image noise, brightness, and skin tone. It can also optimize parameters such as exposure and color temperature of the captured scene. In some embodiments, the ISP can be located within camera 193.

[0075] The camera 193 can be used to capture still images or videos. The object generates an optical image through the lens and projects it onto the photosensitive element. The photosensitive element can be a charge coupled device (CCD) or a complementary metal-oxide-semiconductor (CMOS) phototransistor. The photosensitive element converts the light signal into an electrical signal, and then transmits the electrical signal to the ISP to be converted into a digital image signal. The ISP outputs the digital image signal to the DSP for processing. The DSP converts the digital image signal into an image signal in a standard format such as RGB, YUV, etc.

[0076] In the embodiment of the present application, the mobile phone 100 may include multiple cameras 193, for example, a first camera for implementing a binocular camera, a second camera, and a third camera for implementing a ToF camera. Of course, other cameras may also be included, such as a fourth camera for implementing a structured light camera, and the present application does not impose any specific restrictions on the number and type of cameras.

[0077] The external memory interface 120 can be used to connect an external memory card, such as a Micro SD card, to expand the storage capacity of the mobile phone 100. The external memory card communicates with the processor 110 through the external memory interface 120 to implement data storage functions. For example, files such as music and videos can be stored on the external memory card.

[0078] The internal memory 121 can be used to store computer executable program codes, which include instructions. The processor 110 executes various functional applications and data processing of the mobile phone 100 by running the instructions stored in the internal memory 121. The internal memory 121 may include a program storage area and a data storage area. Among them, the program storage area can store the software code of the operating system and at least one application required for a function (such as a camera application, an album application, a WeChat application). The data storage area can store data created during the use of the mobile phone 100 (such as image data, audio data, a phone book, etc.). The internal memory 121 can be used to store the computer executable program code of the image blur processing method proposed in the embodiment of the present application, which includes instructions. The processor 110 can call the computer executable program code of the image blur processing method stored in the internal memory 121, so that the mobile phone 100 completes the image blur processing solution proposed in the embodiment of the present application.

[0079] The internal memory 121 can also store the image obtained after blurring. Exemplarily, the image after blurring and the original image (i.e., the image before blurring) can be stored correspondingly. Exemplarily, after the mobile phone 100 detects an instruction for opening the original image, it displays the original image, and a mark can be displayed on the original image. When the mark is triggered, the mobile phone 100 opens the blurred image (the image after blurring the original image); or, after the mobile phone 100 detects an instruction for opening the blurred image, it displays the blurred image, and a mark is displayed on the blurred image. When the mark is triggered, the mobile phone 100 opens the original image (the original image corresponding to the blurred image).

[0080] The mobile phone 100 can implement audio functions such as music playback and recording through the audio module 170, the speaker 170A, the receiver 170B, the microphone 170C, the headphone jack 170D, and the application processor.

[0081] Keys 190 include a power button, a volume button, etc. Keys 190 may be mechanical keys or touch keys. Mobile phone 100 may receive key inputs and generate key signal inputs related to user settings and function control of mobile phone 100.

[0082] Motor 191 can generate vibration prompts. Motor 191 can be used for incoming call vibration prompts, and can also be used for touch vibration feedback. For example, touch operations acting on different applications (such as taking pictures, audio playback, etc.) can correspond to different vibration feedback effects. For touch operations acting on different areas of the display screen 194, motor 191 can also correspond to different vibration feedback effects. Different application scenarios (for example: time reminders, receiving messages, alarm clocks, games, etc.) can also correspond to different vibration feedback effects. The touch vibration feedback effect can also support customization.

[0083] The indicator 192 may be an indicator light, which may be used to indicate the charging status, power level changes, messages, missed calls, notifications, etc.

[0084] The SIM card interface 195 is used to connect a SIM card. The SIM card can be connected to and disconnected from the mobile phone 100 by inserting it into or removing it from the SIM card interface 195. The mobile phone 100 can support 1 or N SIM card interfaces, where N is a positive integer greater than 1. The SIM card interface 195 can support Nano SIM cards, Micro SIM cards, SIM cards, and the like. Multiple cards can be inserted into the same SIM card interface 195 at the same time. The types of the multiple cards can be the same or different. The SIM card interface 195 can also be compatible with different types of SIM cards. The SIM card interface 195 can also be compatible with external memory cards. The mobile phone 100 interacts with the network through the SIM card to implement functions such as calls and data communications. In some embodiments, the mobile phone 100 uses an eSIM, i.e., an embedded SIM card. The eSIM card can be embedded in the mobile phone 100 and cannot be separated from the mobile phone 100.

[0085] It should be understood that the structures illustrated in the embodiments of the present application do not constitute a specific limitation on the mobile phone 100. In other embodiments of the present application, the mobile phone 100 may include more or fewer components than shown, or may combine or separate certain components, or have different component arrangements. The illustrated components may be implemented in hardware, software, or a combination of software and hardware.

[0086] The following combination Figure 1 The mobile phone architecture shown in FIG. 1 introduces the process of the image blurring processing method provided by the embodiment of the present application. Figure 2 The figure is a flow chart of the image blurring processing method provided by the embodiment of the present application, which can be applied to Figure 1 The mobile phone 100 shown, or other devices. Figure 1 Taking the mobile phone 100 as an example, the software code of the method can be stored in the internal memory 121, and the processor 110 of the mobile phone 100 runs the software code to implement Figure 2 The image capture process is shown in Figure 2. Figure 2 As shown, the process of the method includes:

[0087] S201: The mobile phone uses a binocular camera and a ToF camera to capture an image of a first scene, obtains an RGB image and a first depth image corresponding to the first scene based on the binocular camera, and obtains depth information of the first scene based on the ToF camera.

[0088] Specifically, the processor 110 in the mobile phone responds to the first instruction to start the binocular camera and the ToF camera to perform a shooting operation. Among them, the first instruction is used to instruct the processor 110 to perform the image shooting operation. The first instruction can be an instruction generated by triggering a function in the application (such as the shooting function in WeChat is started, the camera is started), or it can be an instruction generated based on an input operation performed by the user. The input operation can be a contact operation (such as a click operation or a long press operation) input by the user on the display screen 194, or a somatosensory operation or contactless operation input near the mobile phone, or an operation of recording a voice command (such as "take a photo", "123", etc.). The embodiment of the present application does not make specific restrictions.

[0089] For example, see Figure 3 , the user clicks the shooting button below the shooting screen with his finger on the shooting interface of the camera application. After the display screen 194 detects the touch operation input by the user, the signal corresponding to the detected touch operation is sent to the touch controller in real time. The touch controller converts the signal into touch coordinates and sends them to the processor 110; the processor 110 determines that the touch operation is specifically clicking the "shoot" button on the shooting interface of the camera application based on the received touch coordinates, and the processor 110 responds to the click operation input by the user, turns on the binocular camera and the ToF camera, and performs the shooting operation of the first scene.

[0090] In an embodiment of the present application, the binocular camera includes a first camera and a second camera. When the binocular camera is started under the control of the processor 110, the first camera and the second camera respectively capture a first image of a first scene at a first perspective and a second image at a second perspective, where the first perspective and the second perspective are different (for example, a left perspective and a right perspective, respectively); then, the first camera and the second camera transmit the captured first image and second image to the processor 110, and the processor 110 calculates the first image and the second image based on the binocular parallax principle to obtain a sparse depth map. For the sake of convenience of description, the sparse depth image obtained based on the binocular parallax principle is referred to as the first depth image in this document.

[0091] In a specific implementation, the processor 110 can process the first image and the second image through a hardware module. For example, software code for implementing binocular camera parameter calibration (including intrinsic and extrinsic parameters), binocular image correction (including distortion correction and stereo correction), binocular feature point matching, disparity map or depth map calculation, and virtual viewpoint synthesis using the disparity map or depth map is fixed on a designated chip (e.g., an ISP). By running this software code, the ISP can process the first image and the second image and generate a first depth image, thereby improving the system's computing speed.

[0092] In other embodiments, the binocular camera can be replaced with a multi-camera. Compared to the binocular camera, the multi-camera can include more cameras, such as a first camera, a second camera, and a third camera, where the first camera, the second camera, and the third camera are respectively used to capture images of the first scene from three different perspectives, such as the first image, the second image, and the third image. After obtaining the first image, the second image, and the third image, the processor 110 performs feature point matching and depth calculation on the first image, the second image, and the third image based on the principle of multi-view parallax to obtain a first depth image.

[0093] In an embodiment of the present application, when photographing a first scene, a ToF camera can continuously send light pulses to a target (e.g., each object in the first scene), and then use a sensor to receive the light returned from the object. The processor 110 can obtain the distance of the target object based on the flight (round-trip) time of the light. Specifically, the ToF camera may include a transmitting module and a receiving module, wherein the transmitting module and the receiving module are respectively configured to emit light and receive light reflected from the surface of the photographed object under the control of the processor 110. The processor 110 calculates the distance to each point in the first scene and the mobile phone by calculating the time it takes for the light to be emitted from the ToF camera and returned to the ToF camera, thereby obtaining depth information of the first scene. The transmitting module may be a visible light source, a laser light source (e.g., an infrared light source), etc., and is not specifically limited in the embodiment of the present application. The receiving module may be a camera module, such as a third camera, or the first or second camera of a multiplexed binocular camera. In some embodiments, the depth information of the first scene obtained by the ToF camera may also be expressed as a depth map, such as a third depth map.

[0094] In other embodiments, if the object to be photographed is close to the mobile phone, such as a portrait selfie scene, the ToF camera can be replaced with a structured light camera, that is, the mobile phone can simultaneously use the binocular camera and the structured light camera to capture an image of the first scene, obtain a first depth image of a first resolution based on the binocular parallax of the binocular camera, obtain depth information of the first scene based on the structured light camera, and then use the depth information obtained by the structured light camera to process the first depth image. Among them, the specific implementation method of the structured light camera obtaining the depth information of the first scene includes: after the structured light camera is started by the processor 110, a pre-designed pattern is projected onto the first scene as a reference image (coded light source), and then structured light is projected onto the surface of the object in the first scene, and then the camera is used to receive the structured light pattern reflected from the surface of the object, so that the structured light pattern reflected from the surface of the object obtained by the camera is compared with the reference image, and the depth value of each point in the first scene is calculated by the position and deformation degree of the structured light pattern reflected from the surface of the object, thereby obtaining the depth information of the first scene (or the third depth map).

[0095] S202: The mobile phone densifies the first depth image using the depth information of the first scene acquired by the ToF camera to supplement invalid pixels in the first depth image to obtain a second depth image.

[0096] Specifically, the processor 110 in the mobile phone expands the valid pixels in the invalid depth area around each pixel in the first depth image (that is, the area where the invalid pixels are located, which has no depth value or no accurate depth value) based on the high-precision depth information obtained by ToF.

[0097] For example, see Figure 4A 、 Figure 4B . Figure 4A is the partial image of the first depth image before densification, Figure 4A Positions m2, m3, and m4 around position m1 are invalid pixels with no depth information. Based on the depth information of the first scene acquired by ToF, the depth values at positions m2, m3, and m4 are matched and filled with the corresponding depth values. Similar methods can be used to expand pixels in other invalid pixel areas, but examples are not provided here. Figure 4B is the partial image of the second depth image after densification, Figure 4B and Figure 4A compared to, Figure 4B The number of effective pixels in the image increases, which means the degree of refinement is higher.

[0098] S203: The mobile phone uses the second depth image to perform foreground and background segmentation on the RGB image acquired by the binocular camera to obtain a first foreground image and a first background image.

[0099] The RGB image for foreground and background segmentation may be a left view image or a right view image captured by a binocular camera, or a new RGB image obtained by fusion calculation of the left view image and the right view image, which is not limited here.

[0100] The processor 110 in the mobile phone maps each pixel in the second depth image to the RGB image, so that the depth value corresponding to each pixel in the RGB image can be obtained; then, based on a segmentation threshold, the area in the RGB image with a depth value greater than or equal to the segmentation threshold can be determined as the background area, and the area in the RGB image with a depth value less than the segmentation threshold can be determined as the foreground area, thereby identifying the foreground edge, and then segmenting the foreground and background from the second depth image based on the edge. In some embodiments, the segmentation threshold can be pre-set by the designer and stored in the internal memory 121 of the mobile phone; in other embodiments, the segmentation threshold can also be obtained by the processor 110 by observing the overall depth distribution of the RGB image. Generally speaking, there are usually obvious differences between foreground and background points. The foreground point distribution is generally more concentrated and the depth value is usually much smaller than the background point. Therefore, the position with a larger depth value gradient is generally the edge position of the foreground. For example, see Figure 5A 、 Figure 5B ,in Figure 5A for Figure 3 The RGB image corresponding to the shooting scene shown in FIG, the edge of the foreground of the image, such as Figure 5B shown.

[0101] S204: The mobile phone performs blurring processing on the first background image, and fuses the first foreground image with the blurred first background image to obtain an RGB image with a background blur effect.

[0102] Specifically, the processor 110 in the mobile phone may refine the design of (spatial domain) filter parameters based on the second depth image, and use the filter to perform blurring processing on the first background image.

[0103] In an embodiment of the present application, there may be various types of filters used for image blurring, such as a mean filter, a Gaussian filter (Gaussian blur), a median filter, or a bilateral filter, etc., which are not specifically limited by the embodiment of the present application. The parameters of the filter that requires refined design may include the size of the filter, the weight coefficient corresponding to each pixel in the filter, and the shape of the filter. Among them, the size of the filter and the weight coefficient corresponding to each pixel in the filter determine the degree of blurring of backgrounds at different depths, and can be designed based on depth information, while the shape of the filter affects the shape of the blurred light spot and can be pre-set by a technician.

[0104] For example, take the mean filter as an example. First, a template is given to each pixel in the background image. The template includes the adjacent pixels around the pixel (that is, the 8 pixels around the pixel as the center, forming a filter template, as shown in the following example. Figure 6 Then use all the pixels in the template (i.e. Figure 6The pixel value of the pixel point is replaced by the average value of the pixel of the 9 pixels shown in the figure. In specific implementation, the size of the template can be adjusted, for example, it can also be 5*5, 7*7, etc., and can be adjusted according to the required degree of blurring. This is not limited in the embodiments of this application. One possible design is that the higher the intensity of blurring required, the larger the template.

[0105] In the embodiment of the present application, the area with a larger depth value may have a higher degree of blurring. Figure 5A 、 5B The second depth image shown is Figure 5A 、 5B As shown, the first background image includes image elements corresponding to three trees from left to right, wherein the distance between the tree located behind the left of the portrait and the phone is greater than the distance between the tree located behind the right of the portrait and the phone. Therefore, when blurring the background image, the blurring degree of the image elements corresponding to the tree located behind the left of the portrait can be greater than the blurring degree of the image elements corresponding to the tree located behind the right of the portrait, as shown in FIG. Figure 7 shown.

[0106] In the above scheme, when the binocular camera captures the depth image of the first scene, the ToF camera is controlled to capture the first scene at the same time, the RGB image and the first depth image of the first scene are obtained based on the binocular camera, and the high-precision depth information of the first scene is obtained based on the ToF camera. Then, the first depth image captured by the binocular camera is densified using the high-precision depth information obtained by the ToF camera to obtain a second depth image, and then the second depth image is used to segment the foreground and background of the RGB image. This can improve the accuracy of foreground edge recognition, and even in scenes with similar textures, colors, etc. in the foreground and background, a good edge segmentation effect can be obtained; and in the blurring process, the greater the depth of field, the higher the blurring intensity, so that the blurred background has a good sense of layering, a natural transition, and a blurring effect closer to a real optical defocus effect.

[0107] Furthermore, since the blurring process is performed at a relatively small size, the background portion obtained by the final sampling may have problems such as blur and moiré. In view of this, after obtaining the RGB image with the background blur effect, the mobile phone in the embodiment of the present application can also fuse the RGB image with the background blur effect and the original RGB image (the RGB image before blurring), which can effectively eliminate problems such as foreground blur and moiré.

[0108] The specific implementation method of the fusion can be to perform weighted addition of the pixels in the RGB image with the background blur effect and the pixels in the original RGB image. The weight coefficients corresponding to the pixels at different positions in the scene can be different. For example, the larger the position area corresponding to the depth value, the larger the weight coefficient of the pixel in the RGB image with the background blur effect, and the smaller the weight coefficient of the pixel in the original RGB image. Of course, there can be other fusion methods in the specific implementation, and the embodiments of the present application are not limited thereto.

[0109] In some other embodiments of the present application, the above step S204 may also be to blur the first foreground image, thereby realizing a depth map with a foreground blur effect. Figure 8 , the method flow for blurring the foreground may include:

[0110] S301: The mobile phone uses a binocular camera and a ToF camera to capture an image of a first scene, obtains an RGB image based on the capture by the binocular camera, obtains a first depth image based on binocular parallax of the binocular camera, and obtains depth information of the first scene based on the ToF camera.

[0111] S302: The mobile phone densifies the first depth image using the depth information of the first scene acquired by the ToF camera to obtain a second depth image.

[0112] S303: The mobile phone uses the second depth image to segment the foreground and background of the RGB image to obtain a first foreground image and a first background image.

[0113] The specific implementation of steps S301 to S303 may refer to the specific implementation of steps S201 to S203 above, which will not be repeated here.

[0114] S304: The mobile phone performs blurring processing on the first foreground image, and fuses the blurred first foreground image with the first background image to obtain an RGB image with a foreground blurring effect.

[0115] When the processor 110 in the mobile phone blurs the foreground, the intensity of blurring can be higher in areas with smaller depth of field, giving the blurring a sense of depth. By fusing the blurred foreground and background, a depth image with a foreground blur effect can be obtained. The specific implementation method of blurring the foreground can refer to the implementation method of blurring the background described above and will not be repeated here.

[0116] For example, the above Figure 5A 、 5B The second depth image shown is Figure 9 As shown in the figure, it is the effect of blurring the foreground image. Figure 9The portrait in the foreground is blurred, while the trees in the background remain clear.

[0117] The above solution combines a ToF camera with the depth image of the first scene taken by the binocular camera to obtain high-precision depth information of the first scene. The high-precision depth information obtained by the ToF camera is used to densify the first depth image taken by the binocular camera and segment the foreground and background, thereby improving the accuracy of foreground edge recognition and optimizing the blurred edge effect. Moreover, during the blurring process, the foreground is blurred using a high-precision second depth image, which can make the blurring intensity higher in areas with smaller depth of field, thereby improving the layering of the blurred foreground.

[0118] Furthermore, in an embodiment of the present application, if the subject being photographed in the shooting scene is a subject of a specific type, CNN can also be used to identify the edges of the subject being photographed, segment the background, and further optimize the foreground edge effect of the blurred image.

[0119] For example, see Background Blur. Figure 10 When photographing a subject of a specified type, the image blur processing method in the embodiment of the present application may include:

[0120] S401: The mobile phone uses a binocular camera and a ToF camera to capture an image of a second scene, obtains an RGB image based on the capture by the binocular camera, obtains a first depth image based on binocular parallax of the binocular camera, and obtains depth information of the second scene based on the ToF camera.

[0121] The specific implementation of step S401 can refer to the specific implementation of step 201 above, which will not be repeated here.

[0122] S402: Use the trained CNN model to perform feature extraction on the RGB image to extract the feature image corresponding to the target object.

[0123] Specifically, the processor 110 pre-trains the CNN using several image instances of other photographed subjects of the same type as the target object, obtains a trained CNN model, and stores it in the internal memory 121, wherein the input of the model is an image, and the output is a feature map corresponding to the photographed subject in the image. After obtaining the RGB image, the processor 110 reads the CNN model from the internal memory 121, runs the CNN model, inputs the RGB image into the CNN model, and the CNN model calculates the RGB image and outputs a feature image corresponding to the target object in the RGB image. The target object may be a subject of a specific type, such as a human face. Of course, in specific implementation, it may also be other subjects, such as animals, flowers, etc., and the embodiments of the present application do not make specific restrictions here.

[0124] For example, let's take a portrait. Figure 11A , user A takes a selfie with his mobile phone, and the corresponding second scene is the scene where user A is. Figure 11A It can be seen that in the second scene, in addition to user A himself, there is another pedestrian B (located behind user A to the right). After the processor 110 captures the RGB image corresponding to the second scene based on the binocular camera, the CNN model can be used to perform feature extraction on the RGB image to extract the feature image corresponding to the target object, such as Figure 11B As shown, the portrait feature image of user A is extracted ( Figure 11B The part indicated by the dotted line).

[0125] S403 : Densify the first depth image using the depth information of the second scene acquired by the ToF camera to obtain a second depth image.

[0126] The specific implementation of step S403 can refer to the specific implementation of step 202 above, which will not be repeated here.

[0127] As an optional implementation, when the processor 110 densifies the second depth image using the depth information of the second scene obtained by the ToF camera, it can also combine the feature map extracted by the above-mentioned CNN model to distinguish the foreground and background, and then densify the foreground and background separately, thereby improving the accuracy of the densification.

[0128] S404: Identify a feature image based on the CNN model, segment the foreground and background in the RGB image, and obtain a second foreground image and a second background image.

[0129] Specifically, the processor 110 in the mobile phone uses the feature image obtained in step S402 as the foreground image, then identifies the edge of the foreground image, and segments the RGB image into foreground and background based on the edge to obtain a second foreground image and a second background image.

[0130] As an optional implementation, in order to further improve the accuracy of foreground and background segmentation, the depth information obtained by the ToF camera (or structured light camera) or the above-mentioned densified second depth image can also be combined to segment the foreground and background. Exemplarily, the processor 110 first preliminarily determines the feature image obtained in step S402 as the foreground image, and then identifies the edge of the foreground image, and determines the boundary area of the foreground and background images based on the identified edge (such as the area within a preset range from the edge); then, the processor 110 corrects the pixels in the boundary area of the foreground and background images based on the depth information or the second depth image obtained by the ToF camera (or structured light camera), for example, dividing the pixels in the boundary area with a depth value greater than or equal to a preset value into the background, and dividing the pixels in the boundary area with a depth value less than a preset value into the foreground, etc.; then, the processor 110 re-identifies the edge of the foreground image, and segments the foreground and background of the RGB image based on the re-identified edge to obtain a second foreground image and a second background image.

[0131] S405 , blurring the second background image, and fusing the blurred second background image with the second foreground image to obtain an RGB image with a background blur effect.

[0132] Among them, the specific implementation method of step S405 can refer to the specific implementation method of the above-mentioned step S204, and will not be repeated in the embodiment of this application. Figure 11C For Figure 11B The image shown in the figure is blurred. Figure 11C It can be seen that in the blurred image, only the face of user A is clear, while the body of user A and pedestrian B in the background are blurred, which can achieve the visual effect of highlighting the face of user A.

[0133] Based on the binocular camera and ToF camera, the above-mentioned technology further combines CNN to identify specific types of photographed subjects and intelligently segment the foreground and background to achieve different blurring effects.

[0134] For example, refer to Figures 11A to 11C In the example shown, if the CNN is not used and only the binocular camera and the ToF camera are used for blurring, then user A's body will also be classified as the foreground due to the small depth values of both the face and body, and thus user A's body will not be blurred. However, after the CNN recognition step is combined, the facial feature map can be identified, so the effect of blurring all parts except the portrait can be achieved. Of course, in specific implementations, the portrait in this article can also refer to the entire human body, that is, both the face and body are recognized as the foreground, and this embodiment of the present invention is not limited here.

[0135] The above is an introduction to the technical solution of this application using the example of blurring a single-frame image. During the specific implementation process, the technical solution of the embodiment of this application is also applicable to the scenario of blurring multiple consecutive frames of images (such as videos or animated images).

[0136] If the image captured by the mobile phone is a continuous multi-frame image, the mobile phone can perform the above steps S201 to S204 for each frame of the image, so that the blur effect of each frame of the image can be optimized. For example, taking the video with background blur effect as an example, see Figure 12 The method for blurring a video in an embodiment of the present application includes:

[0137] S501: The mobile phone starts the binocular camera and the ToF camera to shoot video, obtains multiple frames of continuous RGB images based on the shooting of the binocular camera, obtains a first depth image corresponding to each frame of the multiple frames of continuous RGB images based on the binocular parallax of the binocular camera, and obtains depth information corresponding to each frame of the RGB image based on the ToF camera.

[0138] S502. The mobile phone performs densification processing on a first depth image corresponding to each RGB image in the multiple continuous RGB images by using the depth information corresponding to the RGB image obtained by the second camera, so as to supplement invalid pixels in the first depth image corresponding to the RGB image, and obtain a second depth image corresponding to the RGB image.

[0139] S503: The mobile phone uses the second depth image corresponding to each RGB image in the multiple continuous RGB images to perform foreground and background segmentation on the RGB image to obtain a foreground image and a background image corresponding to each RGB image.

[0140] S504: The mobile phone performs blurring processing on the foreground image or the background image corresponding to each frame of the multiple continuous RGB images to obtain an RGB image with a background blur effect or a foreground blur effect corresponding to each frame of the RGB image. The multiple frames of RGB images with the background blur effect or the foreground blur effect form a video with the background blur effect or the foreground blur effect.

[0141] The specific processing method of each frame of RGB image in the above steps S501-S504 can refer to the specific processing method of the RGB image in the above steps S201-S204, and will not be repeated in the implementation of this application.

[0142] Similarly, during the blurring process, the processor 110 in the mobile phone can also use the second depth image corresponding to each frame of the RGB image or the high-precision depth information obtained by the ToF camera to blur the background of the frame of RGB image, so that the greater the depth of field, the higher the blur intensity, so that the blurred background has a good sense of layering. Of course, the processor 110 can also blur the foreground of the video. The specific implementation method of blurring the foreground of each frame of RGB image can refer to the specific implementation method of blurring the foreground of the RGB image in steps S301-S304 above, and will not be repeated in this embodiment of the application.

[0143] On the basis of the video images captured by the binocular camera, the ToF camera is combined to obtain the depth information of the captured scene. The high-precision depth information obtained by the ToF camera is used to densify the first depth image corresponding to each frame of the RGB image captured by the binocular camera and segment the foreground and background, thereby improving the accuracy of video foreground edge recognition. Even in scenes with similar textures, colors, etc. of the foreground and background, a good edge segmentation effect can be obtained; and in the process of blurring each frame of the RGB image, the second depth image (or the high-precision depth information obtained by the ToF camera) is also used to blur the background or foreground of the RGB image, so that the blur has a good sense of layering and a natural transition, which can improve the visual experience.

[0144] As an optional implementation, after densifying each frame of image and before blurring it, the mobile phone can also use the high inter-frame stability of the ToF camera and CNN to perform time-domain smoothing on each frame of image.

[0145] Specifically, the processor 110 in the mobile phone can perform weighted fusion of N adjacent RGB images in the above-mentioned multiple frames of continuous RGB images, where N is a positive integer greater than 1, and the weighted coefficient of each frame of the RGB image in the front and back N adjacent RGB images can be determined based on the depth information corresponding to the RGB image obtained by the ToF camera, or determined based on the foreground and background edge information of the RGB image identified by the CNN, or the fusion coefficient of the RGB image can be determined by considering the depth information corresponding to the RGB image obtained by the ToF camera and the foreground and background edge information of the RGB image identified by the CNN. The embodiment of the present invention does not limit this. Further, in the specific implementation, all RGB images in the video can be divided into multiple groups of RGB images in chronological order, where each group includes N adjacent RGB images, and then the above-mentioned weighted fusion operation is performed on each group of RGB images, thereby achieving time domain smoothing of all RGB images. Accordingly, when the processor 110 subsequently blurs the video, it specifically blurs the RGB image obtained by weighted fusion.

[0146] Since the depth information between adjacent RGB images obtained by ToF has good stability and the edge information between adjacent RGB images detected by CNN has good stability, weighted fusion of adjacent RGB images based on the depth information between adjacent RGB images obtained by ToF and / or the edge information between adjacent RGB images detected by CNN can make the image transition between adjacent RGB images after weighted fusion more natural, thereby improving the problem of background flickering in the RGB images captured by the binocular camera in the prior art due to the poor inter-frame stability of the binocular camera when shooting videos with a blur effect.

[0147] In the embodiments provided above, the methods provided in the embodiments of the present application are described from the perspective of an electronic device (mobile phone 100) as the execution subject. In order to implement the various functions in the methods provided in the embodiments of the present application, the terminal may include a hardware structure and / or a software module, and implement the above functions in the form of a hardware structure, a software module, or a hardware structure plus a software module. Whether one of the above functions is executed in the form of a hardware structure, a software module, or a hardware structure plus a software module depends on the specific application and design constraints of the technical solution.

[0148] Based on the same technical concept, the embodiment of the present application further provides a computer-readable storage medium, wherein the computer-readable storage medium includes a computer program, and when the computer program is run on an electronic device, the electronic device executes the above-mentioned Figure 2 、 Figure 8 、 Figure 10 、 Figure 12 All or part of the steps described in the method embodiments shown.

[0149] Based on the same technical concept, the embodiment of the present application also provides a program product, including instructions, which, when executed on a computer, causes the computer to execute the above-mentioned Figure 2 、 Figure 8 、 Figure 10 、 Figure 12 All or part of the steps described in the method embodiments shown.

[0150] Based on the same technical concept, the embodiment of the present application also provides a circuit system, which can be one or more chips, such as a system on a chip. In some embodiments, the circuit system can be Figure 1The mobile phone 100 or a component of the mobile phone 100 is shown. The circuit system is used to generate a first control signal, the first control signal is used to control the first camera to perform a shooting operation on a first scene and obtain a first depth image corresponding to the first scene; the circuit system is also used to generate a second control signal, the second control signal is used to control the second camera to perform a shooting operation on the first scene and obtain depth information of the first scene.

[0151] The various implementation modes of this application can be combined arbitrarily to achieve different technical effects.

[0152] The above embodiments are merely used to provide a detailed introduction to the technical solutions of the present application. However, the descriptions of the above embodiments are only intended to help understand the methods of the embodiments of the present application and should not be construed as limiting the embodiments of the present application. Any changes or substitutions that can be easily conceived by those skilled in the art should be included within the scope of protection of the embodiments of the present application.

[0153] As used in the above embodiments, the term “when…” may be interpreted to mean “if…” or “after…” or “in response to determining…” or “in response to detecting…”, depending on the context. Similarly, the phrases “upon determining…” or “if (stated condition or event) is detected” may be interpreted to mean “if determining…” or “in response to determining…” or “upon detecting (stated condition or event)” or “in response to detecting (stated condition or event)”, depending on the context.

[0154] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the process or function described in the embodiment of the present application is generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions can be transmitted from one website, computer, server or data center to another website, computer, server or data center via a wired (e.g., coaxial cable, optical fiber, digital subscriber line) or wireless (e.g., infrared, wireless, microwave, etc.) method. The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that includes one or more available media integrations. The available medium can be a magnetic medium (e.g., a floppy disk, a hard disk, a tape), an optical medium (e.g., a DVD), or a semiconductor medium (e.g., a solid-state hard disk).

[0155] For the purpose of explanation, the foregoing description is described with reference to specific embodiments. However, the above exemplary discussion is not intended to be exhaustive, nor is it intended to limit the present application to the precise forms disclosed. In light of the above teachings, many modifications and variations are possible. The embodiments are selected and described in order to fully illustrate the principles of the present application and its practical application, so that others skilled in the art can take full advantage of the present application and various embodiments with various modifications suitable for the specific purposes contemplated.

Claims

1. A method for image blurring, characterized in that: Applied to an electronic device, the electronic device is provided with a first camera and a second camera, wherein the first camera is a binocular camera or a multi-camera camera, and the second camera is a time-of-flight (ToF) camera or a structured light camera, the method comprising: In response to a first instruction, starting the first camera and the second camera to perform a shooting operation on a first scene, the first camera acquiring a red, green, and blue (RGB) image and a first depth image corresponding to the first scene, and the second camera acquiring depth information of the first scene; Densify the first depth image using the depth information of the first scene to supplement invalid pixels in the first depth image to obtain a second depth image, wherein the number of invalid pixels in the second depth image is less than the number of invalid pixels in the first depth image; Using the second depth image to perform foreground and background segmentation on the RGB image to obtain a foreground image and a background image; The background image or the foreground image is blurred to obtain an RGB image with a background or foreground blur effect.

2. The method according to claim 1, wherein Segmenting the RGB image into foreground and background using the second depth image includes: The foreground and background in the RGB image are segmented using the second depth image and a trained convolutional neural network (CNN) model; wherein the input of the trained CNN model is an RGB image in which the photographed subject is a specific type of subject, and the output is the RGB image portion corresponding to the photographed subject in the image.

3. The method according to claim 2, wherein The specific type of subject is a portrait.

4. The method according to any one of claims 1 to 3, wherein Performing blur processing on the background image or the foreground image to obtain an RGB image with a background blur effect or a foreground blur effect includes: Using the second depth image to blur the background image on the RGB image, wherein a pixel area with a larger depth value in the background image corresponds to a higher degree of blur; fusing the foreground image and the blurred background image to obtain an RGB image with a background blur effect; or The foreground image is blurred, wherein the pixel area with a smaller depth value in the foreground image corresponds to a higher degree of blurring; and the background image and the blurred foreground image are fused to obtain an RGB image with a foreground blurring effect.

5. The method according to any one of claims 1 to 3, wherein After obtaining the RGB image with the background or foreground blur effect, the following steps are also included: The RGB image with the background or foreground blur effect is weightedly fused with the RGB image.

6. The method according to any one of claims 1 to 3, wherein: The first camera acquires an RGB image and a first depth image corresponding to the first scene, including: The first camera captures a video corresponding to the first scene, where the video includes multiple frames of continuous RGB images, and obtains a first depth image corresponding to each RGB image in the multiple frames of continuous RGB images; The second camera acquires depth information of the first scene, including: The second camera obtains depth information corresponding to each RGB image in the multiple consecutive RGB images; Densifying the first depth image using the depth information of the first scene to obtain a second depth image includes: For each RGB image in the multiple continuous RGB images, densify a first depth image corresponding to the RGB image using depth information corresponding to the RGB image acquired by the second camera to supplement invalid pixels in the first depth image corresponding to the RGB image, thereby obtaining a second depth image corresponding to the RGB image; Using the second depth image to perform foreground and background segmentation on the RGB image to obtain a foreground image and a background image, including: Using the second depth image corresponding to each RGB image in the multiple continuous RGB images, the frame RGB image is segmented into a foreground image and a background image corresponding to the frame RGB image; Performing blur processing on the background image or the foreground image to obtain an RGB image with a background blur effect or a foreground blur effect includes: A foreground image or a background image corresponding to each frame of the multiple continuous RGB images is blurred to obtain a depth image with a blurred background or foreground effect corresponding to each frame of the RGB image.

7. The method according to claim 6, wherein Before blurring the foreground image or the background image corresponding to each RGB image in the plurality of continuous RGB images, the method further includes: performing weighted fusion on N adjacent RGB images in the plurality of continuous RGB images, where N is a positive integer greater than 1, and a weighting coefficient corresponding to each RGB image in the N adjacent RGB images is determined based on depth information corresponding to the RGB image frame acquired by the second camera and / or foreground and background edge information of the RGB image frame identified by the CNN; The blurring process is performed on the background image or the foreground image, comprising: The background image or foreground image corresponding to the weighted fused RGB image is blurred.

8. The method according to claim 4, wherein The first camera acquires an RGB image and a first depth image corresponding to the first scene, including: The first camera captures a video corresponding to the first scene, where the video includes multiple frames of continuous RGB images, and obtains a first depth image corresponding to each RGB image in the multiple frames of continuous RGB images; The second camera acquires depth information of the first scene, including: The second camera obtains depth information corresponding to each RGB image in the multiple consecutive RGB images; Densifying the first depth image using the depth information of the first scene to obtain a second depth image includes: For each RGB image in the multiple continuous RGB images, densify a first depth image corresponding to the RGB image using depth information corresponding to the RGB image acquired by the second camera to supplement invalid pixels in the first depth image corresponding to the RGB image, thereby obtaining a second depth image corresponding to the RGB image; Using the second depth image to perform foreground and background segmentation on the RGB image to obtain a foreground image and a background image, including: Using the second depth image corresponding to each RGB image in the multiple continuous RGB images, the frame RGB image is segmented into a foreground image and a background image corresponding to the frame RGB image; Performing blur processing on the background image or the foreground image to obtain an RGB image with a background blur effect or a foreground blur effect includes: A foreground image or a background image corresponding to each frame of the multiple continuous RGB images is blurred to obtain a depth image with a blurred background or foreground effect corresponding to each frame of the RGB image.

9. The method according to claim 5, wherein The first camera acquires an RGB image and a first depth image corresponding to the first scene, including: The first camera captures a video corresponding to the first scene, where the video includes multiple frames of continuous RGB images, and obtains a first depth image corresponding to each RGB image in the multiple frames of continuous RGB images; The second camera acquires depth information of the first scene, including: The second camera obtains depth information corresponding to each RGB image in the multiple consecutive RGB images; Densifying the first depth image using the depth information of the first scene to obtain a second depth image includes: For each RGB image in the multiple continuous RGB images, densify a first depth image corresponding to the RGB image using depth information corresponding to the RGB image acquired by the second camera to supplement invalid pixels in the first depth image corresponding to the RGB image, thereby obtaining a second depth image corresponding to the RGB image; Using the second depth image to perform foreground and background segmentation on the RGB image to obtain a foreground image and a background image, including: Using the second depth image corresponding to each RGB image in the multiple continuous RGB images, the frame RGB image is segmented into a foreground image and a background image corresponding to the frame RGB image; Performing blur processing on the background image or the foreground image to obtain an RGB image with a background blur effect or a foreground blur effect includes: A foreground image or a background image corresponding to each frame of the multiple continuous RGB images is blurred to obtain a depth image with a blurred background or foreground effect corresponding to each frame of the RGB image.

10. The method according to claim 8 or 9, characterized in that Before blurring the foreground image or the background image corresponding to each RGB image in the plurality of continuous RGB images, the method further includes: performing weighted fusion on N adjacent RGB images in the plurality of continuous RGB images, where N is a positive integer greater than 1, and a weighting coefficient corresponding to each RGB image in the N adjacent RGB images is determined based on depth information corresponding to the RGB image frame acquired by the second camera and / or foreground and background edge information of the RGB image frame identified by the CNN; The blurring process is performed on the background image or the foreground image, comprising: The background image or foreground image corresponding to the weighted fused RGB image is blurred.

11. An electronic device, characterized in that: The electronic device includes a first camera, a second camera, and at least one processor, wherein the first camera is a binocular camera or a multi-camera, and the second camera is a time-of-flight (ToF) camera or a structured light camera; The processor is configured to, in response to a first instruction, activate the first camera and the second camera to perform a shooting operation on a first scene; The first camera is configured to capture the first scene and obtain a red, green, and blue (RGB) image and a first depth image corresponding to the first scene; The second camera is configured to capture the first scene and obtain depth information of the first scene; The processor is further configured to perform densification processing on the first depth image using the depth information of the first scene to supplement invalid pixels in the first depth image to obtain a second depth image, wherein the number of invalid pixels in the second depth image is less than the number of invalid pixels in the first depth image; Using the second depth image to perform foreground and background segmentation on the RGB image to obtain a foreground image and a background image; The background image or the foreground image is blurred to obtain an RGB image with a background or foreground blur effect.

12. The electronic device according to claim 11, wherein: When the processor uses the second depth image to perform foreground and background segmentation on the RGB image, the processor is specifically configured to: The foreground and background in the RGB image are segmented using the second depth image and a trained convolutional neural network (CNN) model; wherein the input of the trained CNN model is an RGB image in which the photographed subject is a specific type of subject, and the output is the RGB image portion corresponding to the photographed subject in the image.

13. The electronic device according to claim 12, wherein: The specific type of subject is a portrait.

14. The electronic device according to any one of claims 11 to 13, wherein: When the processor performs blurring processing on the background image or the foreground image, the processor is specifically configured to: Using the second depth image to blur the background image on the RGB image, wherein a pixel region with a larger depth value in the background image corresponds to a higher degree of blur; fusing the foreground image with the blurred background image to obtain an RGB image with a background blur effect; or Performing blurring processing on the foreground image, wherein a pixel region with a smaller depth value in the foreground image corresponds to a higher degree of blurring; The background image and the blurred foreground image are fused to obtain an RGB image with a blurred foreground effect.

15. The electronic device according to any one of claims 11 to 13, wherein: The processor is further configured to: After obtaining the RGB image with the background or foreground blur effect, weighted fusion is performed on the RGB image with the background or foreground blur effect and the RGB image.

16. The electronic device according to any one of claims 11 to 13, characterized in that: The first camera is configured to: capture a video corresponding to the first scene, the video comprising a plurality of consecutive RGB image frames, and obtain a first depth image corresponding to each RGB image frame in the plurality of consecutive RGB image frames; The second camera is used to: obtain depth information corresponding to each RGB image in the multiple frames of continuous RGB images; The processor is configured to: perform densification processing on a first depth image corresponding to each RGB image in the plurality of continuous RGB images using depth information corresponding to the RGB image acquired by the second camera, so as to supplement invalid pixels in the first depth image corresponding to the RGB image, thereby obtaining a second depth image corresponding to the RGB image; Using the second depth image corresponding to each RGB image in the multiple continuous RGB images, the frame RGB image is segmented into a foreground image and a background image corresponding to the frame RGB image; A foreground image or a background image corresponding to each frame of the multiple continuous RGB images is blurred to obtain a depth image with a blurred background or foreground effect corresponding to each frame of the RGB image.

17. The electronic device according to claim 16, wherein: The processor is further configured to: Before blurring the foreground image or the background image corresponding to each RGB image frame in the plurality of continuous RGB images, weighted fusion is performed on N adjacent RGB images in the plurality of continuous RGB images, where N is a positive integer greater than 1, and a weighting coefficient corresponding to each RGB image in the N adjacent RGB images is determined based on depth information corresponding to the RGB image frame acquired by the second camera and / or foreground and background edge information of the RGB image frame identified by the CNN; When performing blurring processing on the background image or the foreground image, the processor is specifically configured to: perform blurring processing on the background image or the foreground image corresponding to the weighted fused RGB image.

18. The electronic device according to claim 14, wherein: The first camera is configured to: capture a video corresponding to the first scene, the video comprising a plurality of consecutive RGB image frames, and obtain a first depth image corresponding to each RGB image frame in the plurality of consecutive RGB image frames; The second camera is used to: obtain depth information corresponding to each RGB image in the multiple frames of continuous RGB images; The processor is configured to: perform densification processing on a first depth image corresponding to each RGB image in the plurality of continuous RGB images using depth information corresponding to the RGB image acquired by the second camera, so as to supplement invalid pixels in the first depth image corresponding to the RGB image, thereby obtaining a second depth image corresponding to the RGB image; Using the second depth image corresponding to each RGB image in the multiple continuous RGB images, the frame RGB image is segmented into a foreground image and a background image corresponding to the frame RGB image; A foreground image or a background image corresponding to each frame of the multiple continuous RGB images is blurred to obtain a depth image with a blurred background or foreground effect corresponding to each frame of the RGB image.

19. The electronic device according to claim 15, wherein: The first camera is configured to: capture a video corresponding to the first scene, the video comprising a plurality of consecutive RGB image frames, and obtain a first depth image corresponding to each RGB image frame in the plurality of consecutive RGB image frames; The second camera is used to: obtain depth information corresponding to each RGB image in the multiple frames of continuous RGB images; The processor is configured to: perform densification processing on a first depth image corresponding to each RGB image in the plurality of continuous RGB images using depth information corresponding to the RGB image acquired by the second camera, so as to supplement invalid pixels in the first depth image corresponding to the RGB image, thereby obtaining a second depth image corresponding to the RGB image; Using the second depth image corresponding to each RGB image in the multiple continuous RGB images, the frame RGB image is segmented into a foreground image and a background image corresponding to the frame RGB image; A foreground image or a background image corresponding to each frame of the multiple continuous RGB images is blurred to obtain a depth image with a blurred background or foreground effect corresponding to each frame of the RGB image.

20. The electronic device according to claim 18 or 19, wherein: The processor is further configured to: Before blurring the foreground image or the background image corresponding to each RGB image frame in the plurality of continuous RGB images, weighted fusion is performed on N adjacent RGB images in the plurality of continuous RGB images, where N is a positive integer greater than 1, and a weighting coefficient corresponding to each RGB image in the N adjacent RGB images is determined based on depth information corresponding to the RGB image frame acquired by the second camera and / or foreground and background edge information of the RGB image frame identified by the CNN; When performing blurring processing on the background image or the foreground image, the processor is specifically configured to: perform blurring processing on the background image or the foreground image corresponding to the weighted fused RGB image.

21. An electronic device, characterized in that: The electronic device includes a first camera, a second camera, at least one processor and a memory, wherein the first camera is a binocular camera or a multi-camera, and the second camera is a time-of-flight (ToF) camera or a structured light camera; The memory is used to store one or more computer programs; when the one or more computer programs stored in the memory are executed by the at least one processor, the electronic device is able to implement the method according to any one of claims 1 to 10.

22. A computer-readable storage medium, characterized in that The computer-readable storage medium includes a computer program, and when the computer program is run on an electronic device, the electronic device is caused to perform the method according to any one of claims 1 to 10.

23. A program product, characterized in that The method comprises instructions, which, when executed on a computer, cause the computer to execute the method according to any one of claims 1 to 10.

24. A circuit system, characterized in that: including one or more chips; The circuit system is configured to generate a first control signal, wherein the first control signal is configured to control the first camera to perform a shooting operation on a first scene and obtain a red, green, and blue (RGB) image and a first depth image corresponding to the first scene; The circuit system is further configured to generate a second control signal, wherein the second control signal is configured to control the second camera to perform a photographing operation on the first scene to obtain depth information of the first scene; Wherein, the first camera is a binocular camera or a multi-camera camera, and the second camera is a time-of-flight ToF camera or a structured light camera; The depth information of the first scene is used to perform densification processing on the first depth image to supplement invalid pixels in the first depth image to obtain a second depth image, wherein the number of invalid pixels in the second depth image is less than the number of invalid pixels in the first depth image; The second depth image is used to perform foreground and background segmentation on the RGB image to obtain a foreground image and a background image; The background image or the foreground image is used to perform blurring processing to obtain an RGB image with a background or foreground blurring effect.

Citation Information

Patent Citations

  • Image processing method and device

    CN108154465A

  • Processing method for improving ToF depth image, 3D image imaging method and electronic device

    CN110009672A

  • Color depth image restoration method based on depth adaptive filtering

    CN110097590A

  • Controlling Image Focus in Real-Time Using Gestures and Depth Sensor Data

    US20150022649A1