Image processing method and related apparatus

CN121190335BActive Publication Date: 2026-09-11HONOR DEVICE CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202410766393.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-06-13
Publication Date
2026-09-11
Estimated Expiration
2044-06-13

AI Technical Summary

Technical Problem

[0003]然而,一些应用所提供的人像模式中,对图像的虚化边缘过渡不够平滑,影响图像虚化的显示效果

Benefits of technology

[0050] It should be understood that the second to sixth aspects of this application correspond to the technical solutions of the first aspect of this application, and the beneficial effects achieved by each aspect and the corresponding feasible implementation are similar, and will not be repeated here.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121190335B_ABST
    Figure CN121190335B_ABST
Patent Text Reader

Abstract

The image processing method and related device provided by the embodiments of the present application relate to the technical field of terminals. The method comprises: performing virtualization processing on the foreground region and the mid-background region respectively, including performing multi-scale processing and multi-scale virtualization on the foreground image, and fusing the foreground region after multi-scale virtualization processing and the mid-background region after virtualization processing, so as to obtain an image with virtualization effects of the foreground region and the mid-background region, and improve the virtualization display effect of the image.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of terminal technology, and in particular to image processing methods and related devices. Background Technology

[0002] Some applications on electronic devices support photography functions, such as camera apps, thus meeting users' photography needs. Camera apps offer various shooting modes, including night mode, portrait mode, and video mode. When users take photos in portrait mode, the electronic device can blur the image.

[0003] However, in the portrait mode provided by some applications, the transition of the blurred edges of the image is not smooth enough, which affects the display effect of the image blur. Summary of the Invention

[0004] The image processing method and related apparatus provided in this application embodiment can perform blurring processing on the foreground region and the middle and background regions respectively, including multi-scale processing and multi-scale blurring of the foreground image, and merging the foreground region after multi-scale blurring processing and the middle and background regions after blurring processing, thereby obtaining an image with blurring effect in both the foreground region and the middle and background regions, and improving the blurring display effect of the image.

[0005] In a first aspect, embodiments of this application provide an image processing method, the method comprising:

[0006] In response to a shooting operation, a first image is acquired; information about the foreground and background images of the first image is acquired; the foreground image information is downsampled to obtain information about M foreground images of different resolutions, where M is an integer greater than 1; the information about the M foreground images of different resolutions is blurred separately to obtain information about M blurred foreground images; the information about the M blurred foreground images is fused to obtain a fused foreground image; the information about the background image is blurred to obtain a blurred background image; the fused foreground image and the blurred background image are then fused together to obtain a second image, which is a blurred image. In this way, the foreground and background images are blurred separately, including multi-scale processing and multi-scale blurring of the foreground image, and then the multi-scale blurred foreground image and the blurred background image are fused together to obtain an image where both the foreground and background have a blurred effect.

[0007] One possible implementation involves fusing the fused foreground image and the blurred background image to obtain a second image. This includes: generating fusion weights for each pixel in M ​​foreground images of different resolutions; and upsampling and fusing the M foreground images of different resolutions based on the fusion weights. This upsampling and fusing of the M foreground images of different resolutions based on the fusion weights results in a more natural image fusion and smoother edge transitions.

[0008] One possible implementation involves generating fusion weights for each pixel in M ​​foreground images of different resolutions. This includes: calculating pixel-level fusion weights, where each pixel in the foreground image has its own weight; and optimizing the pixel-level fusion weights using a Markov Random Field (MRF) function to generate the fusion weights. Since pixel-level fusion weights may lack sufficient spatial smoothness, to improve the continuity of the foreground fusion region and reduce fluctuations in fusion weights between pixels, the pixel-level fusion weights can be optimized using a Markov Random Field (MRF) function, thereby obtaining smooth fusion weights at various scales.

[0009] In one possible implementation, the pixel-level fusion weights are related to the structural information, hue, and color of the foreground image. Thus, the pixel-level fusion weights calculated based on the structural information, hue, and color of the foreground image can result in a generated image with better display effects in terms of hue and color.

[0010] In one possible implementation, the pixel-level fusion weight W_foreground satisfies the following formula:

[0011] W_foreground=W_foreground_structure*W_foreground_tone*W_foreground_color.

[0012] Here, W_foreground_structure represents the weights corresponding to the structural information of the foreground image, W_foreground_tone represents the weights corresponding to the hue of the foreground image, and W_foreground_color represents the weights corresponding to the color of the foreground image. W_foreground_structure, W_foreground_tone, and W_foreground_color are all obtained based on the weight generation function Weight_Gen_Func. In this way, the Weight_Gen_Func function can perform Gaussian weighting or smoothing on the gradient information, hue statistics, and color statistics, making the distribution of the generated pixel-level fusion weights in gradient, hue, and color smoother and more continuous.

[0013] In one possible implementation, the MRF function satisfies the following formula:

[0014] MRF_Fusion_Weight_Pyra=MRF(per_pixel_fusion_weight_Pyra).

[0015] Here, `MRF_Fusion_Weight_Pyra` represents the fusion weight of each pixel after optimization by the MRF function, and `per_pixel_fusion_weight_Pyra` represents the weight of each pixel in the M foreground images with different resolutions. `per_pixel_fusion_weight_Pyra` is calculated based on the pixel-level fusion weights. In this way, the MRF function can be used to optimize the weights of each pixel in M ​​foreground images with different resolutions, improving the continuity of the foreground fusion region, reducing fluctuations in fusion weights between pixels, and thus improving the image display effect.

[0016] In one possible implementation, based on fusion weights, upsampling and image fusion are performed on M foreground images of different resolutions. This includes: using a first fusion method, based on fusion weights, fusing the foreground image of layer N and the background image of layer N to obtain the image fusion result of layer N, where N is an integer greater than 1 and less than or equal to M; upsampling the image fusion result of layer N to obtain the upsampled result of layer N; fusing the foreground image of layer N-1 and the background image of layer N-1 to obtain the image fusion result of layer N-1; and fusing the upsampled result of layer N and the image fusion result of layer N-1. In this way, upsampling the image can generate a higher resolution image. Using the first fusion method for multi-scale fusion and reconstruction allows for fusion of different layers of images, i.e., images of different frequency bands, according to different rules, thereby obtaining images with better foreground blurring effects.

[0017] In one possible implementation, the first fusion method satisfies the following formula:

[0018] Fusion_out_layger_N=W_foreground_N*foreground_N+W_background_N*background_N.

[0019] Wherein, Fusion_out_layger_N represents the image fusion result of layer N, foreground_N represents the foreground image of layer N, W_foreground_N represents the fusion weight corresponding to the foreground image of layer N, background_N represents the background image of layer N, and W_background_N represents the fusion weight corresponding to the background image of layer N. In this way, by defining the weights of the low-resolution layer and the current layer, and by using more of the low-resolution foreground layer, a more blurred foreground effect can be obtained.

[0020] In one possible implementation, the fused foreground image and the blurred background image are image-fused to obtain a second image, including: employing a second fusion method to perform image fusion of the fused foreground image and the blurred background image to obtain a second image; the second fusion method satisfies the following formula:

[0021] Fusion_Result=Front_Scene_Bokeh*Front_Scene_Fusion_Mask+Back_Scene_Bokeh*(1-Front_Scene_Fusion_Mask).

[0022] In this model, Fusion_Result is the second image, Front_Scene_Bokeh is the fused foreground image, Back_Scene_Bokeh is the blurred background image, and Front_Scene_Fusion_Mask is the fusion mask, which represents the weight ratio of the fused foreground image in the second image. Thus, the second fusion method can control the fusion ratio of the two images by using transparency. The formula for the second fusion method is relatively simple, easy to implement, and has relatively low computational complexity, making it easier to achieve a smooth transition effect during image fusion.

[0023] In one possible implementation, the first fusion method includes Laplace blending, and the second fusion method includes Alpha blending. Thus, the first fusion method achieves image fusion through multi-scale decomposition and reconstruction, resulting in smooth edge transitions in the merged image and preserving more image details. The second fusion method controls the fusion ratio of the two images using transparency (Alpha value), thereby achieving a smooth transition effect during image fusion.

[0024] In one possible implementation, obtaining information about the foreground and background images of the first image includes: acquiring target information based on the depth map and focus position of the first image. The target information includes information about the foreground and background images. The target information is used to distinguish between the foreground and background images in the first image; the foreground image is the image before the focus position in the first image, and the background image is the image after the focus position in the first image. Thus, based on the foreground image information, the electronic device can perform multi-scale processing to obtain foreground information at different resolutions. Furthermore, the foreground and background image information can be blurred to obtain an image where both the foreground and background images have a blurred effect, improving the image's blurred display effect.

[0025] One possible implementation includes shooting in portrait mode. In portrait mode, blurring both the foreground and background images effectively highlights the subject, making it appear clearer and more three-dimensional, and enhancing the overall sense of space in the image.

[0026] Secondly, embodiments of this application provide an image processing apparatus, which may be an electronic device, a chip or chip system within an electronic device. The apparatus may include a processing unit. The processing unit is used to implement any processing-related method performed by the electronic device in the first aspect or any possible implementation of the first aspect. When the apparatus is an electronic device, the processing unit may be a processor. The apparatus may further include a storage unit, which may be a memory. The storage unit is used to store instructions, and the processing unit executes the instructions stored in the storage unit to cause the electronic device to implement a method described in the first aspect or any possible implementation of the first aspect. When the apparatus is a chip or chip system within an electronic device, the processing unit may be a processor. The processing unit executes the instructions stored in the storage unit to cause the electronic device to implement a method described in the first aspect or any possible implementation of the first aspect. The storage unit may be a storage unit within the chip (e.g., a register, cache, etc.), or a storage unit located outside the chip within the electronic device (e.g., a read-only memory, random access memory, etc.).

[0027] For example, the processing unit is configured to acquire a first image in response to a shooting operation; and to acquire information about a foreground image and a background image of the first image; specifically, it is further configured to downsample the information of the foreground image to obtain information about M foreground images of different resolutions; to perform blurring processing on the information of the M foreground images of different resolutions respectively to obtain information about M blurred foreground images; to perform image fusion on the information of the M blurred foreground images to obtain a fused foreground image; to perform blurring processing on the information of the background image to obtain a blurred background image; and to perform image fusion on the fused foreground image and the blurred background image to obtain a second image.

[0028] In one possible implementation, the processing unit is used to generate fusion weights for each pixel in M ​​foreground images of different resolutions; and is also used to upsample and fuse the M foreground images of different resolutions based on the fusion weights.

[0029] In one possible implementation, the processing unit is used to calculate pixel-level fusion weights and also to optimize the pixel-level fusion weights based on the Markov Random Field (MRF) function to generate fusion weights.

[0030] In one possible implementation, the pixel-level fusion weights are related to the structural information of the foreground image, the hue of the foreground image, and the color of the foreground image.

[0031] In one possible implementation, the pixel-level fusion weight W_foreground satisfies the following formula:

[0032] W_foreground=W_foreground_structure*W_foreground_tone*W_foreground_color.

[0033] Wherein, W_foreground_structure is the weight corresponding to the structural information of the foreground image, W_foreground_tone is the weight corresponding to the hue of the foreground image, and W_foreground_color is the weight corresponding to the color of the foreground image. W_foreground_structure, W_foreground_tone, and W_foreground_color are all obtained based on the weight generation function Weight_Gen_Func.

[0034] In one possible implementation, the processing unit, for the MRF function, satisfies the following formula:

[0035] MRF_Fusion_Weight_Pyra=MRF(per_pixel_fusion_weight_Pyra).

[0036] Wherein, MRF_Fusion_Weight_Pyra is the fusion weight of each pixel after optimization by the MRF function, and per_pixel_fusion_weight_Pyra is the weight of each pixel in the M foreground images with different resolutions. per_pixel_fusion_weight_Pyra is calculated based on the pixel-level fusion weight.

[0037] In one possible implementation, the processing unit is configured to perform image fusion on the foreground image and the background image of the Nth layer using a first fusion method and based on fusion weights to obtain the image fusion result of the Nth layer; it is also configured to upsample the image fusion result of the Nth layer to obtain the upsampled result of the Nth layer; specifically, it is also configured to perform image fusion on the foreground image and the background image of the (N-1)th layer to obtain the image fusion result of the (N-1)th layer; and to perform image fusion on the upsampled result of the Nth layer and the image fusion result of the (N-1)th layer.

[0038] In one possible implementation, the first fusion method satisfies the following formula:

[0039] Fusion_out_layger_N=W_foreground_N*foreground_N+W_background_N*background_N.

[0040] Wherein, Fusion_out_layger_N is the image fusion result of layer N, foreground_N is the foreground image of layer N, W_foreground_N is the fusion weight corresponding to the foreground image of layer N, background_N is the background image of layer N, and W_background_N is the fusion weight corresponding to the background image of layer N.

[0041] In one possible implementation, the processing unit is used to perform image fusion of the fused foreground image and the blurred background image using a second fusion method to obtain a second image.

[0042] In one possible implementation, the first fusion method includes Laplace Blending, and the second fusion method includes Alpha Blending.

[0043] In one possible implementation, the processing unit is used to obtain target information based on the depth map and focus position of the first image.

[0044] In one possible implementation, the shooting operation includes shooting operations in portrait mode scenes.

[0045] Thirdly, embodiments of this application provide an electronic device including one or more processors and a memory, the memory being coupled to one or more processors, the memory being used to store computer program code, the computer program code including computer instructions, and one or more processors being used to invoke the computer instructions to perform the method described in the first aspect or any possible implementation of the first aspect.

[0046] Fourthly, embodiments of this application provide a computer-readable storage medium storing a computer program or instructions that, when executed on a computer, cause the computer to perform the methods described in the first aspect or any possible implementation thereof.

[0047] Fifthly, embodiments of this application provide a computer program product including a computer program, which, when run on a computer, causes the computer to perform the methods described in the first aspect or any possible implementation thereof.

[0048] Sixthly, this application provides a chip or chip system including at least one processor and a communication interface. The communication interface and the at least one processor are interconnected via a circuit. The at least one processor is used to run computer programs or instructions to perform the methods described in the first aspect or any possible implementation of the first aspect. The communication interface in the chip can be an input / output interface, pins, or circuits, etc.

[0049] In one possible implementation, the chip or chip system described above in this application further includes at least one memory storing instructions. The memory can be an internal storage unit of the chip, such as a register or cache, or it can be a storage unit of the chip itself (e.g., read-only memory, random access memory, etc.).

[0050] It should be understood that the second to sixth aspects of this application correspond to the technical solutions of the first aspect of this application, and the beneficial effects achieved by each aspect and the corresponding feasible implementation are similar, and will not be repeated here. Attached Figure Description

[0051] Figure 1 A schematic diagram of a shooting scene provided for an embodiment of this application;

[0052] Figure 2 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application;

[0053] Figure 3 A schematic diagram of the software structure of an electronic device provided in an embodiment of this application;

[0054] Figure 4 A schematic diagram illustrating the use of mid-background blurring to blur the foreground, provided as an embodiment of this application;

[0055] Figure 5 A flowchart illustrating an image processing method provided in an embodiment of this application;

[0056] Figure 6 A flowchart of dual-camera depth calculation provided in this application embodiment;

[0057] Figure 7 A schematic diagram of depth range provided for an embodiment of this application;

[0058] Figure 8 This is a schematic diagram illustrating the acquisition of foreground mask information provided in an embodiment of this application;

[0059] Figure 9 An illustration of the effect of multi-scale processing of foreground information provided in an embodiment of this application;

[0060] Figure 10 An example of multi-scale blurred foreground information provided in this application embodiment;

[0061] Figure 11 This is a schematic diagram illustrating the acquisition of mid-background mask information provided in an embodiment of this application;

[0062] Figure 12 A schematic diagram illustrating the calculation process of foreground fusion mask information provided in an embodiment of this application;

[0063] Figure 13 A schematic diagram of a Laplace fusion method provided in an embodiment of this application;

[0064] Figure 14 This is a schematic diagram of an image fusion process provided in an embodiment of this application;

[0065] Figure 15 A schematic diagram illustrating an image processing method provided in an embodiment of this application;

[0066] Figure 16 This is a schematic diagram of the structure of a chip provided in an embodiment of this application. Detailed Implementation

[0067] To facilitate a clear description of the technical solutions in the embodiments of this application, some terms and technologies involved in the embodiments of this application will be briefly introduced below:

[0068] 1. Depth of field (DOF): refers to the range of the image that is in focus before and after the image is focused when it is captured.

[0069] Depth of field can be affected by factors such as focal length, aperture, and shooting distance. For example, the longer the focal length, the shallower the depth of field, and the shorter the focal length, the deeper the depth of field; the larger the aperture, the shallower the depth of field, and the smaller the aperture, the deeper the depth of field; the closer to the object, the shallower the depth of field, and the farther away, the deeper the depth of field.

[0070] 2. Circle of Confusion (CoC): Describes the degree of blurriness at a point in an image. Ideally, a point light source, when imaged by a camera, forms a single point on the imaging plane. However, due to limitations of focal length and / or camera hardware, a point light source will form a circle with a certain diameter on the imaging plane; this circle can be called the circle of confusion. In some embodiments, the circle of confusion may also be referred to as the circle of confusion size.

[0071] For example, such as Figure 1As shown in Figure a, objects 1, 2, and 3 are the objects being photographed. When a camera captures an image, it can have different aperture sizes, which can be represented by aperture opening. For example, the distance between 101 and 102 in the figure can represent the aperture opening.

[0072] The distance from the camera to the object being photographed is called the object distance. The object distance from the camera to each object can be the same or different. Figure 1 In scenario 'a', the distance from the camera to object 1 can be less than the distance from the camera to object 2, and the distance from the camera to object 2 can be less than the distance from the camera to object 3. It can be understood that when the camera focuses on object 2, the location of object 2 can be called the focal point or focusing position. The focal point can be understood as the point or area where the camera is focused, and it is usually the clearest part of the image. In some scenarios, object 1 can be used as the foreground or background, object 2 as the midground, and object 3 as the background or background.

[0073] The distance from the camera to the imaging plane can be called the image distance; in some scenarios, the imaging plane can also be called the focal plane. The camera can project images of object 1, object 2, and object 3 onto the imaging plane, thereby generating their respective images. For example, Figure 1 b can be the imaging effect on the imaging plane, which can include the circle of confusion of object 1, the circle of confusion of object 2, and the circle of confusion of object 3.

[0074] 3. Terminology

[0075] In the embodiments of this application, terms such as "first" and "second" are used to distinguish identical or similar items with substantially the same function and purpose. For example, "first chip" and "second chip" are used only to distinguish different chips and do not limit their order of execution. Those skilled in the art will understand that terms such as "first" and "second" do not limit the quantity or execution order, and that "first" and "second" do not necessarily imply that they are different.

[0076] It should be noted that, in the embodiments of this application, the terms "exemplary" or "for example" are used to indicate examples, illustrations, or descriptions. Any embodiment or design scheme described as "exemplary" or "for example" in this application should not be construed as being more preferred or advantageous than other embodiments or design schemes. Specifically, the use of terms such as "exemplary" or "for example" is intended to present the relevant concepts in a specific manner.

[0077] In this application embodiment, "at least one" refers to one or more, and "more than one" refers to two or more. "And / or" describes the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A alone, A and B simultaneously, or B alone, where A and B can be singular or plural. The character " / " generally indicates that the preceding and following related objects are in an "or" relationship. "At least one of the following" or similar expressions refer to any combination of these items, including any combination of single or plural items. For example, at least one of a, b, or c can represent: a, b, c, ab, a--c, bc, or abc, where a, b, and c can be single or multiple.

[0078] 4. Electronic equipment

[0079] The electronic devices in this application embodiment can also be any form of terminal device. For example, electronic devices may include: mobile phones, tablet computers, handheld computers, laptops, mobile internet devices (MIDs), wearable devices, virtual reality (VR) devices, augmented reality (AR) devices, wireless terminals in industrial control, wireless terminals in self-driving, wireless terminals in remote medical surgery, wireless terminals in smart grids, wireless terminals in transportation safety, wireless terminals in smart cities, wireless terminals in smart homes, cellular phones, cordless phones, session initiation protocol (SIP) phones, wireless local loop (WLL) stations, personal digital assistants (PDAs), handheld devices with wireless communication capabilities, computing devices or other processing devices connected to a wireless modem, in-vehicle devices, wearable devices, electronic devices in 5G networks, or future evolved public land mobile communication networks (PLANs). The embodiments of this application do not limit the scope of electronic devices in a mobile network (PLMN).

[0080] By way of example and not limitation, in this embodiment, the electronic device can also be a wearable device. Wearable devices, also known as wearable smart devices, are a general term for devices that utilize wearable technology to intelligently design and develop everyday wearables, such as glasses, gloves, watches, clothing, and shoes. Wearable devices are portable devices that are worn directly on the body or integrated into the user's clothing or accessories. Wearable devices are not merely hardware devices, but also achieve powerful functions through software support, data interaction, and cloud interaction. Broadly speaking, wearable smart devices include those that are feature-rich, large in size, and can achieve complete or partial functions without relying on a smartphone, such as smartwatches or smart glasses, as well as those that focus on a specific type of application function and require the use of other devices such as smartphones, such as various smart bracelets and smart jewelry for vital sign monitoring.

[0081] Furthermore, in this application embodiment, the electronic device can also be an electronic device in the Internet of Things (IoT) system. IoT is an important part of the future development of information technology. Its main technical feature is to connect objects to the network through communication technology, thereby realizing an intelligent network of human-machine interconnection and object-to-object interconnection.

[0082] The electronic equipment in the embodiments of this application may also be referred to as: user equipment (UE), mobile station (MS), mobile terminal (MT), access terminal, user unit, user station, mobile station, mobile station, remote station, remote terminal, mobile device, user terminal, terminal, wireless communication equipment, user agent, or user device, etc.

[0083] In this embodiment, the electronic device or various network devices include a hardware layer, an operating system layer running on top of the hardware layer, and an application layer running on top of the operating system layer. The hardware layer includes hardware such as a central processing unit (CPU), a memory management unit (MMU), and memory (also called main memory). The operating system can be any one or more computer operating systems that implement business processing through processes, such as Linux, Unix, Android, iOS, or Windows. The application layer includes applications such as browsers, address books, word processing software, and instant messaging software.

[0084] For example, Figure 2 A schematic diagram of the electronic device is shown.

[0085] The electronic device may include a processor 110, an external memory interface 120, an internal memory 121, a universal serial bus (USB) interface 130, a charging management module 140, a power management module 141, a battery 142, an antenna 1, an antenna 2, a mobile communication module 150, a wireless communication module 160, an audio module 170, a speaker 170A, a receiver 170B, a microphone 170C, a headphone jack 170D, a sensor module 180, buttons 190, a motor 191, an indicator 192, a camera 193, a display screen 194, and a subscriber identification module (SIM) card interface 195, etc. The sensor module 180 may include a pressure sensor 180A, a gyroscope sensor 180B, a barometric pressure sensor 180C, a magnetic sensor 180D, an accelerometer sensor 180E, a distance sensor 180F, a proximity sensor 180G, a fingerprint sensor 180H, a temperature sensor 180J, a touch sensor 180K, an ambient light sensor 180L, a bone conduction sensor 180M, etc.

[0086] It is understood that the structures illustrated in the embodiments of this application do not constitute a specific limitation on the electronic device. In other embodiments of this application, the electronic device may include more or fewer components than illustrated, or combine some components, or split some components, or have different component arrangements. The illustrated components may include hardware, software, or a combination of software and hardware.

[0087] Processor 110 may include one or more processing units, such as application processors (APs), modem processors, graphics processing units (GPUs), image signal processors (ISPs), controllers, video codecs, digital signal processors (DSPs), baseband processors, and / or neural network processing units (NPUs). These different processing units may be independent devices or integrated into one or more processors. The controller can generate operation control signals based on instruction opcodes and timing signals to control instruction fetching and execution.

[0088] The processor 110 may also include a memory for storing instructions and data. In some embodiments, the memory in the processor 110 is a cache memory. This memory can store instructions or data that the processor 110 has just used or that are used repeatedly. If the processor 110 needs to use the instruction or data again, it can directly retrieve it from the aforementioned memory. This avoids repeated accesses, reduces the waiting time of the processor 110, and thus improves the efficiency of the system. For example, in the embodiments of this application, the processor 110 can be used to process foreground or mid-background mask information, multi-scale foreground information, and blurring of foreground or mid-background, etc.

[0089] Internal memory 121 can be used to store computer executable program code, including instructions. Internal memory 121 may include a program storage area and a data storage area. The program storage area may store the operating system, at least one application program required for a function, etc. The data storage area may store data created during the use of the electronic device, etc. Furthermore, internal memory 121 may include high-speed random access memory, and may also include non-volatile memory, such as at least one disk storage device, flash memory device, universal flash storage (UFS), etc. Processor 110 executes various functional applications and data processing of the electronic device by running instructions stored in internal memory 121 and / or instructions stored in memory disposed in the processor. For example, in this embodiment, internal memory 121 may be used to store code related to autofocus algorithms, depth-of-field calculation, and foreground or mid-background blurring, etc.

[0090] The electronic device implements display functions through a GPU, a display screen 194, and an application processor. The GPU is a microprocessor for image processing, connected to the display screen 194 and the application processor. The GPU performs mathematical and geometric calculations and is used for graphics rendering. The processor 110 may include one or more GPUs, which execute program instructions to generate or modify display information. For example, in this embodiment, the electronic device can implement shooting functions through an ISP, a camera 193, a video codec, a GPU, a display screen 194, and an application processor.

[0091] In some embodiments, the electronic device may include one or N cameras 193, where N is a positive integer greater than 1. The cameras 193 may be used to capture still images or videos.

[0092] Display screen 194 is used to display images, videos, etc. Display screen 194 includes a display panel. In some embodiments, an electronic device may include one or N displays screens 194, where N is a positive integer greater than 1.

[0093] Figure 3This is a software structure block diagram of an electronic device according to an embodiment of this application. The layered architecture divides the software into several layers, each with a clear role and division of labor. Layers communicate with each other through software interfaces. In some embodiments, the Android system and hardware are layered, from top to bottom as the application layer, application framework layer, hardware adaptation layer (HAL), driver layer, and hardware layer.

[0094] The application layer, also known as the application layer, can include a series of application packages. For example... Figure 3 As shown, the application package can include applications such as camera, gallery, and video. Applications can include system applications and third-party applications.

[0095] The application framework layer, also known as the application framework layer or framework layer, provides application programming interfaces (APIs) and programming frameworks for applications in the application layer. The framework layer can include some predefined functions.

[0096] like Figure 3 As shown, the Framework layer can include a camera access interface. This interface includes camera management and camera device functionality. Camera applications in the application layer can call this interface to manage camera devices, etc.

[0097] The application layer and framework layer run in a virtual machine. The virtual machine executes the Java files of the application layer and framework layer as binary files. The virtual machine is used to perform functions such as object lifecycle management, stack management, thread management, security and exception management, and garbage collection. For example, in the embodiments of this application, the virtual machine can be used to implement functions such as dual-camera depth calculation, autofocus, and image fusion.

[0098] The hardware abstraction layer (HAL) can be a wrapper around hardware drivers, providing a unified interface for upper-layer applications to call them. The HAL includes a camera hardware abstraction layer and a camera algorithm library. The camera hardware abstraction layer can include multiple camera devices. The camera algorithm library includes modules for dual-camera depth calculation, autofocus (AF) algorithms, depth-of-field calculation, foreground mask acquisition, foreground information multi-scale processing, foreground multi-scale blurring, foreground multi-scale fusion, background mask acquisition, background blurring, foreground fusion mask calculation, and foreground and mid-background fusion.

[0099] The dual-camera depth calculation module can perform dual-camera depth calculation based on the main road image and the secondary road image to obtain enhanced depth information (depth map).

[0100] The autofocus algorithm module can be used to automatically adjust the position of the camera by analyzing the sharpness or contrast of the image to obtain focus position information.

[0101] The depth-of-field calculation module can calculate the depth-of-field range by combining the aperture opening of the camera.

[0102] The foreground mask acquisition module can combine depth information and foreground depth range, and filter out the area corresponding to the foreground depth range from the depth information as the foreground area mask, thereby obtaining foreground mask information.

[0103] The foreground information multi-scale processing module can combine the original image with foreground mask information to select foreground region image information. After selecting the foreground region image information, the foreground information multi-scale processing module can also perform multi-scale processing on the foreground region image information to obtain foreground information schematic diagrams at different resolutions.

[0104] The foreground multi-scale blur module can be used to perform blurring processes of different intensities and types at various scales.

[0105] The foreground multi-scale blending module can be used to adjust the blending ratio of the blurring effect at each scale and the blending method at multiple scales, thereby obtaining a better multi-scale blended foreground area.

[0106] The background mask acquisition module can combine depth information, mid-field depth of field range, and background depth of field range to filter out areas from the depth information that correspond to the mid- and background depth of field ranges, and use these areas as masks for the mid- and background regions, thereby obtaining mid- and background mask information.

[0107] The background blur module can blur the mid- and background areas. For example, the mid- and background blur module can use blur algorithms to blur the mid- and background areas.

[0108] The foreground fusion mask calculation module can perform downsampling, mask smoothing, and upsampling on the foreground mask information to obtain the foreground fusion mask information.

[0109] The foreground and midground background fusion module can perform image fusion on the foreground and midground background areas.

[0110] The driver layer is used to drive hardware resources. It can include multiple driver modules, such as camera device drivers, DSP drivers, and GPU drivers. For example, in this embodiment, the camera algorithm library can be used to send digital signals to the DSP driver, enabling the DSP driver to call the DSP in the hardware layer for digital signal processing. The DSP can then return the processed digital signals to the camera algorithm library via the DSP driver. The camera algorithm library is also used to send digital signals to the GPU driver, enabling the GPU driver to call the GPU in the hardware layer for digital signal processing. The GPU can then return the processed image data to the camera algorithm library via the GPU driver.

[0111] The hardware layer includes image sensors, image signal processors (ISPs), digital signal processors (DSPs), and graphics processing units (GPUs). The image sensor can output images based on its capture mode, such as portrait mode. The ISP can perform corresponding processing on the images generated by the image sensor.

[0112] It should be noted that the embodiments of this application are only illustrated using the Android system. In other operating systems (such as Windows system, iOS system, etc.), as long as the functions implemented by each functional module are similar to those in the embodiments of this application, the solution of this application can also be implemented.

[0113] Some applications on electronic devices support photography functions, such as camera apps, thus meeting users' photography needs. When users take pictures using camera apps, the apps can provide various shooting modes, including night mode, portrait mode, and video mode.

[0114] When a user takes a photo in portrait mode, the electronic device can blur the image. This blurring process can include blurring the foreground area of ​​the image and / or blurring the mid-ground and background areas of the image.

[0115] In some embodiments, the foreground region can also be referred to as shallow depth of field, and blurring the foreground region of an image can be simply referred to as foreground blurring or foreground depth blurring. The mid-ground and background regions can also be referred to as background regions, and blurring the mid-ground and background regions of an image can be simply referred to as mid-ground blurring, mid-ground depth blurring, or background blurring. For ease of description, the following descriptions will use foreground region, mid-ground and background region, foreground blurring, and mid-ground blurring.

[0116] However, in some applications' portrait modes, one possible implementation is to blur the mid- and background areas of the image, without blurring the foreground area in the generated image. Alternatively, another possible implementation is to blur the foreground area using the same method as blurring the mid- and background areas, resulting in an image blurring effect that does not meet expectations.

[0117] For example, Figure 4 This image illustrates the result of blurring the foreground area using a mid-to-background blurring technique in portrait mode. Specifically, Figure 4 'a' represents the original image, which can be understood as an image without blurring. For example, the original image may include a YUV format image or an RGB format image. In this original image, the foreground area may include trees 401, the midground area may include people 402, and the background area may include hills 403 and the sun 404. In some embodiments, the original image may also be referred to as the background image. After blurring the foreground area using a midground / background blurring method, as shown... Figure 4 As shown in b, the trees 401 in the foreground area are blurred.

[0118] As can be seen from the image, the transition edge between the foreground and mid-ground / background areas is not smooth, but rather exhibits an unnatural or abrupt change. This phenomenon can be described as a binary cutoff at the image edge. This occurs because the pixel values ​​in the foreground and mid-ground / background areas may suddenly jump from one extreme to another, such as changing from 0 to 255 or vice versa, resulting in an unnatural edge transition and affecting the image display.

[0119] In view of this, the image processing method provided in this application embodiment can perform blurring processing on the foreground region and the middle and background regions respectively, including multi-scale processing and multi-scale blurring of the foreground image, and merging the foreground region after multi-scale blurring processing and the middle and background regions after blurring processing, thereby obtaining an image with blurring effect in both the foreground region and the middle and background regions, and improving the blurring display effect of the image.

[0120] Figure 5 A flowchart illustrating an image processing method according to an embodiment of this application is shown.

[0121] The image processing method of this application embodiment may include (1) blurring of the foreground region, (2) blurring of the middle and background regions, and (3) image fusion of the foreground region and the middle and background regions.

[0122] (1) Blurring the foreground area.

[0123] In this embodiment, the electronic device can combine depth information and focus information to obtain foreground region mask information and complete multi-scale foreground blurring processing. The foreground blurring process may include (1.1) calculation of dual-camera depth, (1.2) acquisition of focus position, (1.3) calculation of depth of field range using simulated aperture settings, (1.4) acquisition of foreground mask information, (1.5) multi-scale processing of foreground information, (1.6) multi-scale blurring of foreground information, and (1.7) multi-scale fusion of foreground information.

[0124] (1.1) Calculation of depth of dual cameras.

[0125] Dual cameras can be understood as two cameras, which can include a main camera and a secondary camera, also referred to as main and secondary cameras.

[0126] It is understandable that the dual cameras of electronic devices have already been aligned during factory calibration and online calibration before the dual-camera depth calculation is performed.

[0127] Factory-calibrated dual-camera alignment can be understood as the factory performing preliminary calibration of the dual cameras during the production process, ensuring accurate initial alignment and parameter settings. This allows the electronic equipment to apply the factory-calibrated distortion parameters and the primary and secondary camera homography matrices to the image and perform online calibration.

[0128] Distortion parameters can be used to correct geometric distortion in cameras, thereby improving the accuracy of depth calculations with dual cameras. It's understandable that when an electronic device uses two cameras to simultaneously capture the same scene, the images will differ slightly due to the different positions of the main and secondary cameras. The Homography matrix can convert the image captured by the secondary camera into an image captured by the main camera, thus aligning the images and reducing differences between them.

[0129] Online calibration of dual-camera alignment can be understood as dynamic calibration of the dual cameras during actual use. Online calibration can reduce the impact of environmental changes, camera position changes, and other factors, resulting in more accurate alignment of the dual cameras. Online calibration yields aligned image pairs, which can include a primary image (left_img) and a secondary image (right_img). Left_img can be understood as the image captured by the primary camera, and right_img as the image captured by the secondary camera.

[0130] like Figure 6As shown, the dual-camera depth calculation module of the electronic device can perform dual-camera depth calculation based on left_img and right_img to obtain enhanced depth information (depth map). In some embodiments, this depth information can also be called a depth map or parallax information, which can represent the positional difference information of the same object in two images.

[0131] On one hand, the dual-camera depth calculation module can use AI segmentation algorithms to segment the left_img image, thereby obtaining portrait segmentation information (portrait segmentation map). In some scenarios, this portrait segmentation information can also be called a portrait segmentation map or a segmentation map.

[0132] Artificial intelligence segmentation algorithms can train models on image data containing human figures. The training model can include convolutional neural networks (CNN), fully convolutional networks (FCN), segmentation models (DeepLab), etc., and this application embodiment does not limit the specific model.

[0133] In a possible implementation, model training may include feature extraction from the image, such as extracting edges, textures, and / or shapes; model training may also involve feature fusion, such as fusing features from different levels to capture more detailed information; model training may also use a loss function to measure the segmentation performance of the trained model and perform optimization. Furthermore, the trained model can output portrait segmentation information, which may be the same size as the input image, and the value of each pixel in the portrait segmentation information can indicate whether the pixel represents a person or the background.

[0134] In this way, the dual-camera depth calculation module can identify and separate the human figure from the background in an image based on artificial intelligence segmentation algorithms, thereby obtaining human figure segmentation information.

[0135] On the other hand, the dual-camera depth calculation module can also use left_img and right_img to perform dual-camera depth calculation (stereo_depth_calc) to obtain depth information.

[0136] In a possible implementation, the dual-camera depth calculation module can also perform feature matching on `left_img` and `right_img`, finding identical feature points in both images. These identical feature points could include the edges or corners of an object in the image. For each pair of identical feature points, the dual-camera depth calculation module can calculate the positional difference of that feature point in the two images. Based on the depth information and the known distance between the cameras, methods such as triangulation can be used to calculate the depth of each feature point, thus obtaining the depth information. It's understandable that the closer the object is to the camera, the greater the relative depth information; conversely, the farther the object is from the camera, the smaller the relative depth information.

[0137] After acquiring the portrait segmentation information and depth information, the dual-camera depth calculation module can fuse the portrait segmentation information and depth information through a depth map enhancement algorithm to obtain the depth information after portrait edge enhancement.

[0138] In a possible implementation, the dual-camera depth calculation module can use edge detection algorithms, such as the Sobel algorithm and the Canny algorithm, to perform edge detection on the depth information and generate a depth edge map. The dual-camera depth calculation module can also use edge detection algorithms to perform edge detection on the portrait segmentation information and generate a segmentation edge map. Furthermore, the dual-camera depth calculation module can fuse the edge information from the depth edge map and the segmentation edge map, for example, by using a logical OR weighted average to fuse the edge information, thereby generating a fused edge map.

[0139] Furthermore, the dual-camera depth calculation module can also enhance the edges of the fused edge map. For example, the dual-camera depth calculation module can add depth gradients to the edge regions to make the edges more distinct; and / or, the dual-camera depth calculation module can perform interpolation in the edge regions to make the edges smoother; and / or, the dual-camera depth calculation module can correct the depth information of the edge regions based on the portrait segmentation information to make the edges more consistent with the true contours of the portrait, etc. The specific method used by the dual-camera depth calculation module to enhance the edges of the fused edge map is not limited in the embodiments of this application.

[0140] (1.2) Obtaining the focus position.

[0141] The autofocus algorithm module of an electronic device can acquire focus position information. In some scenarios, this focus position information can also be referred to as focus position information or focus subject depth information.

[0142] In a possible implementation, the AF algorithm module can analyze the image's sharpness or contrast and automatically adjust the camera's position to obtain the autofocus code (AF Code) or the converted actual distance information of the focus position. The AF Code is a numerical value or code generated when calculating the focus position. The AF Code can represent the specific position where the camera needs to move, thereby achieving a better focusing effect.

[0143] (1.3) Simulate aperture settings to calculate depth of field range.

[0144] In a possible implementation, using the focused position as a reference, the depth-of-field calculation module of the electronic device can calculate the depth-of-field range in conjunction with the aperture opening of the camera. It can be understood that a larger aperture opening results in a shallower depth of field, and a smaller aperture opening results in a deeper depth of field.

[0145] For example, the formula for calculating depth of field (DOF) can satisfy the following formula:

[0146]

[0147] Where C is the circle of confusion size, s is the object distance, D is the aperture opening, f is the focal length, and N is the aperture size. N can be calculated by dividing the focal length f by the aperture opening D, i.e., N = f / D.

[0148] Understandably, when calculating depth of field (DOF), the depth of field range calculation module can obtain parameters such as the circle of confusion size C, aperture opening D, aperture size N, and focal length f from the camera module. After focusing, the autofocus module in the electronic device can calculate the object distance s, and then the depth of field range calculation module can obtain the object distance s based on the autofocus module.

[0149] like Figure 7 As shown, after calculating the depth of field (DOF), the depth of field range calculation module can combine the distance from the camera to the focus position (Focus_pos) to calculate the foreground depth of field range, midground depth of field range, and background depth of field range.

[0150] The foreground depth of field range can be the range from the camera position to the focus position (Focus_pos) minus half the depth of field (DOF). If the camera position is 0, then the foreground depth of field range is [0, Focus_pos - DOF / 2]. In some embodiments, the foreground depth of field range can also be referred to as the foreground blurred area. It is understood that 0 can also be a position near the aperture or the camera, etc., and this application embodiment does not limit this.

[0151] The mid-range depth of field can be defined as the area from the focus position (Focus_pos) minus half the depth of field (DOF) to the focus position (Focus_pos) plus half the depth of field (DOF), i.e., the mid-range depth of field is [Focus_pos - DOF / 2, Focus_pos + DOF / 2]. The focus position (Focus_pos minus half the depth of field) can be understood as the closest distance from the camera to the mid-range depth of field; the focus position (Focus_pos plus half the depth of field) can be understood as the farthest distance from the camera to the mid-range depth of field. In some embodiments, the mid-range depth of field can also be understood as the area of ​​sharp image.

[0152] The depth of field range for the background can be from the focus point (Focus_pos) plus half the depth of field (DOF) to infinity. Alternatively, it can be understood as the distance from the camera to the furthest point in the mid-field depth of field, extending to infinity; that is, the background depth of field range is [Focus_pos + DOF / 2, Infinity]. In some embodiments, the background depth of field range can also be referred to as the background blur area or the background blur area.

[0153] (1.4) Obtaining foreground mask information.

[0154] After calculating the foreground depth range, the foreground mask acquisition module of the electronic device can combine the depth information and the foreground depth range, and filter the area corresponding to the foreground depth range from the depth information as the foreground area mask, thereby obtaining the foreground mask information.

[0155] Figure 8 A schematic diagram illustrating the acquisition of foreground mask information is shown.

[0156] in, Figure 8 'a' is the original image, in which the foreground region may include trees 801, the midground region may include people 802, and the background region may include hills 803 and the sun 804.

[0157] Figure 8 'b' represents a schematic diagram of the depth information of the original image. The foreground, midground, and background regions each correspond to different depth information. Figure 8 In b, the foreground region may include first depth information, which includes the depth information corresponding to trees; the midground region may include second depth information, which includes the depth information corresponding to people; and the background region may include third depth information, which includes the depth information corresponding to hills and the depth information corresponding to the sun.

[0158] Figure 8'c' is a schematic diagram of the acquired foreground mask information. Based on the depth information and the foreground depth range, regions corresponding to the foreground depth range can be selected from the depth information as the foreground mask. For example, Figure 8 The light-colored or white areas in c, such as the areas corresponding to trees, can be represented as foreground mask information; Figure 8 The dark or black areas in c can correspond to the mid-background mask information.

[0159] The foreground mask information can be a binary image or matrix of the same size as the original image. The foreground mask information can be used to distinguish between foreground and mid-ground / background regions in the image. In this foreground mask information, the value of each pixel can be 0 or 1. For example, in the foreground mask information, a pixel with a value of 1 can represent the foreground region, and a pixel with a value of 0 can represent the mid-ground / background region. Optionally, the pixel value can also be represented by true or false. For example, in the foreground mask information, a pixel with a value of true can represent the foreground region, and a pixel with a value of false can represent the mid-ground / background region. The specific representation of pixel values ​​is not limited in this embodiment.

[0160] Optionally, the foreground mask acquisition module can also perform filtering and / or morphological operations on the foreground mask information, such as dilation and erosion, to obtain better foreground mask information.

[0161] (1.5) Multi-scale processing of foreground information.

[0162] After obtaining the foreground mask information, the foreground information multi-scale processing module of the electronic device can combine the original image and use the foreground mask information to select the image information of the foreground region. After selecting the image information of the foreground region, the foreground information multi-scale processing module can perform multi-scale processing on the image information of the foreground region to obtain foreground information schematic diagrams at different resolutions.

[0163] For example, the foreground information multi-scale processing module can obtain foreground information at multiple scales by downsampling, such as 2-3 scales.

[0164] like Figure 9 As shown, this example demonstrates how to obtain foreground information at three scales. For instance, Figure 9 'a' can be a schematic diagram of foreground information of the same size as the original image. Figure 9 'b' can be a schematic diagram of foreground information at 1 / 4 size of the original image. Figure 9 c can be a schematic diagram of foreground information at 1 / 16 size of the original image.

[0165] In this embodiment, the resolution of the same size as the original image can be understood as a high resolution, and the corresponding image can be simply referred to as a high-resolution image; the resolution of 1 / 4 the size of the original image and the resolution of 1 / 16 the size of the original image can be understood as low resolution, and the corresponding image can be simply referred to as a low-resolution image. For example, Figure 9 'a' can be understood as a schematic diagram of the foreground information corresponding to the high-resolution image. Figure 9 The 'b' can be understood as a schematic diagram of the foreground information corresponding to the low-resolution image. It is understood that a high-resolution image can also include resolutions similar in size to the original image, and a low-resolution image can also include resolutions near 1 / 4 the size of the original image or resolutions smaller than 1 / 4 the size of the original image.

[0166] Optionally, the multi-scale can also include other numbers of scales. It is understood that if the multi-scale consists of two scales, for example, a scale the same size as the original image and a scale one-quarter the size of the original image, the computational load on the electronic device is relatively small. However, due to the reduced number of blurring scales, the foreground blurring effect is poor, and edge transitions may not be smooth enough. If multiple scales are selected, such as 3-6 scales, the computational load on the electronic device is relatively large. However, due to the increased number of blurring scales, the foreground blurring effect is better, and edge transitions are smoother. Therefore, depending on the actual situation of different electronic devices, a compromise can be made in selecting the number of multi-scales. The specific number of scales selected for processing is not limited in this embodiment.

[0167] (1.6) Multi-scale blurring of foreground information.

[0168] After acquiring multi-scale foreground information, the foreground multi-scale blurring module of the electronic device can perform blurring processing of different intensities and types at each scale. For example, it can perform blurring processing of different intensities and types on foreground information corresponding to high-resolution images and foreground information corresponding to low-resolution images.

[0169] Figure 10 The diagram shows the effect of blurring the foreground information at three scales.

[0170] For example, Figure 10 'a' can represent a foreground infographic of the same scale as the original image. Figure 10 b can represent a pair of pairs ... Figure 10 'a' is subjected to the first level of blurring. Figure 10 c can represent the pair of Figure 10 The 'a' is subjected to a second level of blurring, and the degree of blurring in the second level is greater than that in the first level.

[0171] Figure 10 The 'd' can represent a foreground infographic of 1 / 4 size from the original image. Figure 10 The 'e' can represent the pair of 'e's'. Figure 10 The d is subjected to a third level of blurring. Figure 10 f can represent the pair of Figure 10 The value of d is blurred with a fourth intensity, and the degree of blurring of the fourth intensity is greater than that of the third intensity. The third intensity and the first intensity, or the third intensity and the second intensity, can be the same or different; the fourth intensity and the first intensity, or the fourth intensity and the second intensity, can be the same or different.

[0172] Figure 10 The 'g' can represent a foreground infographic of 1 / 16th the size of the original image. Figure 10 h can represent the pair of Figure 10 The g is subjected to a fifth level of blurring. Figure 10 The i can represent the pair of pairs ... Figure 10 The value of g is subjected to a sixth level of blurring, and the degree of blurring at the sixth level is greater than that at the fifth level. The fifth and third levels, or the fifth and fourth levels, can be the same or different; the sixth and third levels, or the sixth and fourth levels, can also be the same or different.

[0173] Understandably, the type and intensity of blur at various scales can be adjusted by electronic devices, enabling foreground blur adjustment to have the ability to adjust in various frequency bands, thereby improving the edge blur effect and providing freedom for texture areas.

[0174] (1.7) Multi-scale fusion of the foreground region.

[0175] It is understandable that the foreground multi-scale fusion module of an electronic device may be affected by a variety of factors in the blurring effect of an image. For example, factors affecting the blurring effect may include one or more of the following: space-varying blur, non-circular blur, cat-eye bokeh effect towards corners, aperture blade shapes, clipping, aberrations (aberratios or bokeh fringing), non-uniform bokeh intensity at large aperture settings, diffraction, etc.

[0176] Space-varying can refer to the fact that the degree of blurring of an image varies at different spatial locations, such as the difference in blurring between the center and the edges of the image.

[0177] Non-circular blur and cat-eye bokeh effects towards corners refer to bokeh effects at the edges or corners of an image that are not circular but rather resemble a cat's eye shape. Bokeh can represent the blurred effect of parts of an image outside the focal point. For example, good bokeh can soften and round the background light spots, thus highlighting the subject and making the image more visually appealing.

[0178] Aperture blade shapes represent the shape and number of aperture blades in a camera. The shape and number of aperture blades can affect the bokeh effect, thus affecting the shape and quality of the blurred areas in the image. For example, a larger number of rounder aperture blades can produce a smoother and more rounded bokeh, while a smaller number of straighter aperture blades may produce polygonal bokeh.

[0179] Clipping can refer to the irregular shape of the bokeh that may occur when light passes through the aperture when the aperture blades are not fully open or closed.

[0180] Aberratios can represent various aberrations in an optical system, including spherical aberration, chromatic aberration, coma, etc. These aberrations can cause image distortion, blurring, or color shift. For example, aberrations can include bokeh fringing, which manifests as color shift or halos at the edges of the bokeh.

[0181] Non-uniform bokeh intensity at large aperture settings indicates that the brightness and intensity of the bokeh are uneven when the aperture is wide.

[0182] Diffraction is an optical phenomenon in which light bends and spreads as it passes through a slit or around an obstacle.

[0183] To reduce the impact of the above factors on the blurring effect, in this embodiment, the foreground multi-scale fusion module can adjust the blurring effect fusion ratio and multi-scale fusion method at each scale, thereby obtaining a better multi-scale fused foreground region.

[0184] (2) Blurring of the mid-background area.

[0185] In this embodiment of the application, the electronic device can combine depth information and focus information to obtain the mask of the mid-field area and / or the mask of the background area, and complete the blurring process of the mid-field and / or the blurring process of the background.

[0186] Understandably, if the focus point is in the mid-field area, it means that the mid-field area needs to be clearly displayed and does not need to be blurred. In this case, the electronic device can obtain the mask of the background area and blur the background, without needing to obtain the mask of the mid-field area or blur the mid-field.

[0187] If the focus point is in the foreground area, it means that the mid-ground area needs to be blurred. The electronic device can then acquire the mask of the mid-ground area and / or the mask of the background area, and complete the blurring of the mid-ground and / or the background.

[0188] The following example, with the focus point in the mid-field region, illustrates the process by which an electronic device acquires the mask of the background region and performs background blurring. The background blurring process may include (2.1) acquiring the background mask information and (2.2) blurring the background information.

[0189] (2.1) Obtaining background mask information.

[0190] After calculating the background depth range, the background mask acquisition module of the electronic device can combine the depth information and the background depth range, and select the area corresponding to the background depth range from the depth information as the mask of the background area, thereby obtaining the background mask information.

[0191] Figure 11 A schematic diagram illustrating the acquisition of background mask information is shown.

[0192] in, Figure 11 'a' is the original image, in which the foreground region may include trees 1101, the midground region may include people 1102, and the background region may include hills 1103 and the sun 1104.

[0193] Figure 11 'b' represents a schematic diagram of the depth information of the original image. The foreground, midground, and background regions each correspond to different depth information. Figure 11 In b, the foreground region may include first depth information, which includes the depth information corresponding to trees; the midground region may include second depth information, which includes the depth information corresponding to people; and the background region may include third depth information, which includes the depth information corresponding to hills and the depth information corresponding to the sun.

[0194] Figure 11 'c' is a schematic diagram of the acquired background mask information. Based on the depth information and the background depth range, regions corresponding to the background depth range can be selected from the depth information as the background mask. For example, Figure 11 Dark or black areas in 'c', such as the foreground area including the area corresponding to trees, can be corresponding to foreground mask information, and the midground area including the area corresponding to people can be corresponding to midground mask information. Figure 11 The light-colored or white areas in c, such as the areas corresponding to hills and the sun, can be represented as background mask information.

[0195] (2.2) Blurring of background information.

[0196] Background blurring modules in electronic devices can blur the background area. For example, a background blurring module can use blurring algorithms to blur the background area. Blur algorithms can include Gaussian blur and bokeh effect. Gaussian blur makes the background area appear blurred by applying a Gaussian weighted average to the pixels in the background area. Bokeh blur can simulate the bokeh effect of a large-aperture camera, making background light spots appear blurred in a specific shape. For specific details on blurring the background area, please refer to relevant technologies; they will not be elaborated upon here.

[0197] It is understood that in the embodiments of this application, the blurring process of (1) the foreground region and the blurring process of (2) the background region can be executed in any order.

[0198] (3) Image fusion of foreground and mid-ground regions.

[0199] In this embodiment, the electronic device can calculate foreground fusion mask information to achieve the fusion of foreground blurring result and mid-background blurring effect, thereby obtaining an image with blurring effect in both the foreground and mid-background regions. The image fusion process of the foreground region and the mid-background region may include (3.1) calculation of foreground fusion mask information and (3.2) fusion of foreground and mid-background.

[0200] (3.1) Calculate foreground fusion mask information.

[0201] It is understandable that the blurring effect of the foreground differs from that of the background blurring effect in terms of processing method and visual experience. Foreground blurring needs to provide a soft, hazy, and veiled feeling to highlight the information of the subject. In this embodiment, this effect can be achieved by using foreground blending mask information before blending the foreground and mid-ground areas.

[0202] like Figure 12 As shown, foreground fusion mask information can originate from foreground mask information. Based on the foreground mask information, the foreground fusion mask calculation module of the electronic device can perform downsampling, mask smoothing, upsampling, and other processing on the foreground mask information to obtain the foreground fusion mask information.

[0203] Downsampling can reduce the resolution of foreground mask information by decreasing the number of data points or pixels. Downsampling methods can include nearest-neighbor interpolation, bilinear interpolation, and bicubic interpolation, etc., and are not limited to any particular method described in this application.

[0204] Mask smoothing can reduce noise in an image and make image edges softer. Mask smoothing methods can include Gaussian blur, average blur, median filtering, bilateral filtering, etc., and the embodiments of this application are not limited to these.

[0205] Upsampling can preserve and enhance details. Upsampling methods can include bilinear interpolation, bicubic interpolation, or convolutional neural networks, etc., and are not limited to these methods in the embodiments of this application.

[0206] (3.2) Integration of foreground and background.

[0207] In a possible implementation, the foreground and background blending module of the electronic device can employ alpha blending and / or laplace blending to fuse the foreground and background regions. In some embodiments, laplace blending can also be referred to as pyramid blending.

[0208] In one possible implementation, the foreground and midground blending module can use Alpha Blending to merge the foreground and midground regions.

[0209] For example, the foreground and mid-ground background fusion module can read data from the foreground and mid-ground background images, including data such as color and alpha channels, and normalize the alpha values ​​from the 0-255 range to the 0-1 range. Furthermore, the foreground and mid-ground background fusion module can use alpha blending to calculate the fused image result.

[0210] Alpha Blending can be calculated using the following formula:

[0211] Fusion_Result=Front_Scene_Bokeh*Front_Scene_Fusion_Mask+Back_Scene_Bokeh*(1-Front_Scene_Fusion_Mask).

[0212] Wherein, Fusion_Result is the result of image fusion, Front_Scene_Bokeh is the foreground image with a blurred effect, Back_Scene_Bokeh is the mid-background image with a blurred effect, and Front_Scene_Fusion_Mask is the foreground fusion mask, which is a weight matrix used to determine the weight ratio of the foreground image and the mid-background image in the final fused image. For example, the value of the foreground fusion mask can be between 0 and 1.

[0213] The Alpha Blending method allows you to control the blending ratio of two images by using transparency (Alpha value), thus achieving a smooth transition effect when merging images.

[0214] In another possible implementation, the foreground and midground blending module can use the Laplace Blending method to blend the foreground and midground regions.

[0215] Understandably, in the Laplace Blending method, the Laplacian operator can extract high-frequency information from the image; in the Laplacian pyramid, the higher the layer, the higher the frequency. The foreground and midground blending module can perform alpha fusion on the same layer of the Laplacian pyramid. Alpha fusion utilizes the alpha channel to achieve image fusion. The alpha channel represents information about transparency in the image. The alpha value can be calculated from the foreground fusion mask information using methods such as Gaussian blur. The foreground and midground blending module can perform fusion according to different rules on different layers of images, i.e., images of different frequency bands.

[0216] Taking Gaussian blur as an example, images with larger scales on the Gaussian pyramid can contain more high-frequency features, so a smaller Gaussian kernel can be used; images with smaller scales on the Gaussian pyramid can contain fewer high-frequency features, so a smaller Gaussian kernel can be used.

[0217] For example, taking the fusion of image 1 and image 2 as an example, the following steps are performed: Figure 13 As shown, the LaplaceBlending fusion method can include the following steps:

[0218] a. Create a Laplacian pyramid L1 with a certain number of layers, for example, N layers, for image 1, and a Laplacian pyramid L2 with a certain number of layers, for example, N layers, for image 2. The higher the number of layers in the Laplacian pyramid, the better the fusion effect. The number of layers N can be used as a parameter, and N can be an integer greater than 1.

[0219] b. Input a mask that represents the location for image fusion. For example, when fusing images 1 and 2, the left half of the mask image is 255 and the right half is 0, or the left half of the mask image is 0 and the right half is 255. A Gaussian pyramid can be built based on this mask image for subsequent image fusion.

[0220] Image fusion can satisfy the following formula:

[0221] leftImageWeight*leftImage+rightImageWeight*rightImage=outputImage.

[0222] Wherein, leftImage represents the pixel value matrix of the left image in the Gaussian pyramid, leftImageWeight represents the weight coefficient of the left image, which can also be understood as the proportion of the left image in the fusion result, rightImage represents the pixel value matrix of the right image in the Gaussian pyramid, rightImageWeight represents the weight coefficient of the right image, which can also be understood as the proportion of the right image in the fusion result, and outputImage represents the pixel value matrix of the fused output image.

[0223] It is understandable that the values ​​of leftImageWeight and rightImageWeight are both between 0 and 1, and the sum of leftImageWeight and rightImageWeight is 1.

[0224] c. Merge each layer of the Gaussian pyramid.

[0225] The fusion of each layer can satisfy the following formula:

[0226]

[0227] in, This is the Laplacian pyramid fusion result after fusing the i-th layer of image 1 and the i-th layer of image 2. For the i-th layer of image 1, R i The weights corresponding to the i-th layer of image 1 are: For the i-th layer of image 2, (1-R) i) represents the weight corresponding to the i-th layer of image 1, where i is less than or equal to N.

[0228] d. The final image can be reconstructed based on the new pyramid.

[0229] Understandably, the reconstruction process is similar to that of a typical Laplace pyramid. First, the image of the first layer is upsampled and then added to the top layer of the new pyramid to obtain the second layer image. The second layer image is then upsampled and added to the next layer to obtain the third layer image. This process is repeated, and the final result is the result of the Laplace Blending algorithm.

[0230] Using the Laplace Blending method, image fusion can be achieved through multi-scale decomposition and reconstruction, which allows for smooth transitions at the edges of the fused image and preserves more image details.

[0231] Figure 14 This diagram illustrates the image fusion process performed by the foreground and midground background fusion module.

[0232] For example, the image fusion process may include (1) construction of multi-scale gradients, (2) calculation of pixel-level fusion weights, (3) construction of energy functions, (4) optimization of energy functions and generation of weights, (5) single-layer fusion, and (6) upsampling and multi-layer fusion.

[0233] (1) Construction of multi-scale gradients.

[0234] In a possible implementation, the foreground and midground background fusion module can be based on Poisson fusion using multi-scale gradient information. The weight information is optimized and calculated using Markov random field (MRF) functions at multiple scales to perform image fusion.

[0235] Poisson fusion can be used for image editing and compositing. By solving the Poisson equation, Poisson fusion can achieve seamless image fusion, allowing the inserted image region to smoothly transition with the target image at the boundary. It can also achieve a semi-transparent gradient optical blurring effect at multiple scales in the foreground region, thereby reducing obvious stitching artifacts, while ensuring that the mid-ground and background regions have gradient information, texture information, and / or satisfy structural consistency.

[0236] The foreground and midground fusion module can perform multi-scale downsampling of the image and construct gradient information at each layer. This module can also calculate image gradients to detect edge and texture features in the image. The calculation of the image gradient can satisfy the following formula:

[0237] Image_Pyra=Pyra_Down(Input_image).

[0238] Grad_Info_foreground_Pyra=Image_Grad(foreground_Pyra,Grad_Calculator_XY).

[0239] Grad_Info_background_Pyra=Image_Grad(background_Pyra,Grad_Calculator_XY).

[0240] The Pyra_Down function is used to reduce the resolution of an image. Input_image is the input image, which can be the original image, and Image_Pyra is the output image. It can be understood that electronic devices can output multiple images with the same or different resolutions based on this formula; for example, they can output images at 1 / 4 the size of the original image and images at 1 / 16 the size of the original image.

[0241] The `Image_Grad` function is used to calculate image gradients. `foreground_Pyra` is the input image of the foreground region, which can be an image processed by a pyramid. `background_Pyra` is the input image of the mid-ground and background regions, which can also be an image processed by a pyramid. `Grad_Calculator_XY` is the operator used to calculate the gradient, which can include the Sobel operator, etc. `Grad_Info_foreground_Pyra` contains the gradient information of the foreground image pyramid, which can include edge and structural features of the foreground image.

[0242] Understandably, in some scenarios, `foreground_Pyra` can be `Image_Pyra`. This means that an electronic device can first reduce the resolution of the foreground image to obtain the output image `Image_Pyra`, and then use this output image `Image_Pyra` as the input image to calculate the gradient of the foreground image. Similarly, in some scenarios, `background_Pyra` can be `Image_Pyra`. This means that an electronic device can first reduce the resolution of the mid-ground and background images to obtain the output image `Image_Pyra`, and then use this output image `Image_Pyra` as the input image to calculate the gradient of the mid-ground and background images.

[0243] (2) Calculation of pixel-level fusion weights.

[0244] In some embodiments, pixel-level fusion weights can also be referred to as basic fusion weights. The fusion weights W_foreground of the foreground image can satisfy the following formula:

[0245] W_foreground=W_foreground_structure*W_foreground_tone*W_foreground_color.

[0246] Where W_foreground_structure represents the weights related to structural information, and W_foreground_structure tends to select pixels in the foreground image with weaker structural information. W_foreground_structure can satisfy the following formula:

[0247] W_foreground_structure=Weight_Gen_Func(Gaussian_distribution_func*Grad_Info_foreground_Pyra).

[0248] The `Weight_Gen_Func` function is a weight generation function that can be used to Gaussian-weight or smooth gradient information to generate a smoother, more continuous weight distribution. `Gaussian_distribution_func` is a Gaussian distribution function that can also be used to weight or smooth gradient information. `Grad_Info_foreground_Pyra` can be used to represent the gradient information of the foreground image pyramid, which can include edge and structural features of the foreground image.

[0249] W_foreground_tone is a weight related to hue, and it tends to select pixels with low hue variation. W_foreground_tone can satisfy the following formula:

[0250] W_foreground_tone=Weight_Gen_Func(Gaussian_distribution_func*Tone_Statistic(foreground_image)).

[0251] The `Weight_Gen_Func` function is a weight generation function that can be used to Gaussian-weight or smooth tonal statistics to generate a smoother, more continuous weight distribution. `Gaussian_distribution_func` is a Gaussian distribution function that can also be used to weight or smooth tonal statistics. `Tone_Statistic(foreground_image)` can be used to calculate the tonal statistics of the foreground image, which can include the image's brightness distribution, grayscale histogram, etc.

[0252] W_foreground_color is a color-related weight, which tends to select pixels with rich colors. W_foreground_color can satisfy the following formula:

[0253] W_foreground_color=Weight_Gen_Func(Gaussian_distribution_func*Color_Statistic(foreground_image)).

[0254] The `Weight_Gen_Func` function is a weight generation function that can be used to Gaussian-weight or smooth color statistics to generate a smoother, more continuous weight distribution. `Gaussian_distribution_func` is a Gaussian distribution function that can also be used to weight or smooth color statistics. `Color_Statistic(foreground_image)` can be used to calculate the color statistics of the foreground image, which can include the image's color distribution, hue histogram, saturation distribution, etc.

[0255] The fusion weight W_mid_background of the mid-background image can satisfy the following formula:

[0256] W_mid_background=1024-W_foreground.

[0257] (3) Construction of energy function.

[0258] Since pixel-level fusion weights may suffer from insufficient spatial domain smoothness, this application provides a weight update scheme based on energy function optimization, namely a Markov random field weight optimization scheme, to improve the continuity of the foreground fusion region and reduce fluctuations in inter-pixel fusion weights. This scheme satisfies the following formula:

[0259] MRF_Fusion_Weight_Pyra=MRF(per_pixel_fusion_weight_Pyra).

[0260] Here, the MRF function stands for Markov Random Field function, which can be used to optimize the input weight pyramid. MRF_Fusion_Weight_Pyra is the fusion weight pyramid optimized by the MRF, representing the fusion weight for each pixel after considering spatial relationships. per_pixel_fusion_weight_Pyra is the fusion weight pyramid for each pixel, which can be a multi-scale weight matrix used to represent the weight of each pixel at different resolutions.

[0261] It's understandable that Per_Pixel_Fusion_Weight_Pyra can be derived from W_foreground_Pyra information. Electronic devices can execute the Pyra_Down function on the fusion weights W_foreground of the foreground image to obtain the fusion weights W_foreground_Pyra information of the foreground image after resolution reduction. Alternatively, it can be understood as multi-scalening the fusion weights W_foreground of the foreground image to obtain the W_foreground_Pyra information.

[0262] After modeling, the foreground and midground background fusion module can use the graph cut algorithm to optimize the energy function, thereby obtaining smooth fusion weight coefficients at various scales.

[0263] The energy function required for modeling can be constructed as a mapping from pixel-level weights to the field variable Pos. The energy function E(Pos) corresponding to the modeling can satisfy the following formula:

[0264] E(Pos)=Energy_Data(Pos)+lamda*Energy_Smooth(Pos).

[0265] Where Pos is a random value of the field variable representing the position of the fused coordinate point, which can change with the current coordinate value (i,j). Energy_Data(Pos) is the data energy term, Energy_Smooth(Pos) is the smoothing energy term, and lambda is the weight parameter of the smoothing energy term. Lambda can be used to balance the influence of data energy and smoothing energy.

[0266] Energy_Data(Pos) can satisfy the following conditions:

[0267] Energy_Data(Pos)=Per_Pixel_Fusion_Weight if Pos(i,j)=1024.

[0268] Per_Pixel_Fusion_Weight is the fusion weight for each pixel. Pos(i,j) = 1024 is a condition, which means that when the pixel value Pos(i,j) at coordinate (i,j) equals 1024, the data energy Energy_Data(Pos) at that coordinate value is equal to the fusion weight Per_Pixel_Fusion_Weight for each pixel. It is understood that the electronic device can adjust the value of 1024 in the condition according to actual needs, and this application embodiment does not limit this.

[0269] Alternatively, Energy_Data(Pos) can satisfy the following conditions:

[0270] Energy_Data(Pos)=1024-Per_Pixel_Fusion_Weight if Pos(i,j)=0.

[0271] The condition Pos(i,j) = 0 indicates that when the pixel value Pos(i,j) at coordinate (i,j) is equal to 0, the data energy Energy_Data(Pos) at that coordinate is equal to (1024 - Per_Pixel_Fusion_Weight). It is understood that the electronic device can adjust the value of 1024 in the condition according to actual needs, and this application embodiment does not impose any limitations.

[0272] Energy_Smooth(Pos) can satisfy the following formula:

[0273] Energy_Smooth(Pos)=1024-Diff(Pos(i,j),Pos(m,n)).

[0274] Energy_Smooth(Pos) can represent the Pos value measure between the current position (i,j) and the neighboring positions (m,n), where the closer the Pos values ​​are, the smaller the Diff value is.

[0275] (4) Optimization of energy function and generation of weights.

[0276] After obtaining Pos(i,j), the foreground and background fusion module can obtain the value of Pos(i,j) at each pixel position through the energy function E(Pos), thereby obtaining the fusion weight information at each pixel position. Specific optimization methods can include graph cut and other methods.

[0277] (5) Single-layer fusion.

[0278] The single-layer fusion can be referred to the relevant description in (3.2) above on the fusion of foreground and midground background, and will not be repeated here.

[0279] (6) Upsampling and multi-scale fusion.

[0280] After the foreground and midground background fusion module completes the weight generation at each layer across multiple scales, it can perform multi-scale fusion and reconstruction. Multi-scale fusion and reconstruction can satisfy the following formula:

[0281] Fusion_out_layger_N=W_foreground_N*foreground_N+W_background_N*background_N.

[0282] Fusion_out_layger_N-1_Up=upscale(Fusion_out_layger_N).

[0283] Fusion_out_layger_N-1_Fusion=W_foreground_N-1*foreground_N-1+(1024–W_foreground_N-1)*background_N-1.

[0284] Fusion_out_layger_N-1=fusion(Fusion_out_layger_N-1_Up, Fusion_out_layger_N-1_Fusion).

[0285] Wherein, foreground_N is the foreground image to be fused in the Nth layer, W_foreground_N is the weight matrix of the foreground image in the Nth layer, where each element represents the weight of the corresponding pixel, background_N is the mid-background image to be fused in the Nth layer, W_background_N is the weight matrix of the mid-background image in the Nth layer, and W_background_N can be 1024-W_foreground_N, and Fusion_out_layger_N is the result of fusing the foreground image and the mid-background image in the Nth layer.

[0286] Understandably, in some scenarios, if it is not necessary to blur the mid-range image, the mid-to-background image can also be represented as the background image.

[0287] The Upscale function is used to upsample an image to generate a higher resolution image. Fusion_out_layger_N-1_Up is the result obtained by upsampling Fusion_out_layger_N.

[0288] Fusion_out_layger_N-1_Fusion is the result of fusing the foreground image and the mid-ground and background images of layer N-1. foreground_N-1 is the foreground image to be fused in layer N-1, W_foreground_N-1 is the weight matrix of the foreground image in layer N-1, background_N-1 is the mid-ground and background images to be fused in layer N-1, and (1024-W_foreground_N-1) is the weight matrix of the mid-ground and background images in layer N-1.

[0289] Fusion_out_layger_N-1 is the output of the image fusion between the (N-1)th layer and the low-resolution layer. The Fusion function is used to perform image fusion processes, such as weighted averaging, stitching, and convolution. The weights of the low-resolution layer and the current layer are defined by weights, so that more of the low-resolution foreground can be used to obtain a more blurred foreground effect.

[0290] It is understood that in video blurring scenarios, the image processing method of this application embodiment can also be used to blur the foreground area of ​​the video, thereby obtaining a video with blurring effects in both the foreground and middle and background areas.

[0291] The methods of this application will be described in detail below through specific embodiments. The following embodiments can be combined with each other or implemented independently, and the same or similar concepts or processes may not be described again in some embodiments.

[0292] Figure 15 An image processing method according to an embodiment of this application is illustrated. The method includes:

[0293] S1501, in response to the shooting operation, acquire the first image.

[0294] In this embodiment of the application, the shooting operation can be used to obtain a blurred image. The shooting operation may include touch operation, voice control, gesture operation, etc. Touch operation includes, for example, clicking the shooting button. This embodiment of the application does not limit the operation.

[0295] The first image can be understood as an image that has not undergone blurring processing, or as the original image in the above embodiments. The first image may include an image in YUV format or an image in RGB format.

[0296] S1502. Obtain information about the foreground image and the background image of the first image.

[0297] In this embodiment of the application, the information of the foreground image may include foreground mask information.

[0298] In a possible implementation, the electronic device can combine the depth information of the first image with the foreground depth range, and filter out the region corresponding to the foreground depth range from the depth information to obtain the foreground image information. The specific process of obtaining the foreground image information of the first image can be referred to... Figure 5 The relevant descriptions of (1.1) the calculation of dual-camera depth, (1.2) the acquisition of the focus position, (1.3) the calculation of the depth of field range by simulating aperture settings, and (1.4) the acquisition of foreground mask information in the corresponding embodiments will not be repeated here.

[0299] Background image information may include background mask information or mid-background mask information.

[0300] Understandably, if the focus point is in the mid-ground of the first image, it means the mid-ground area needs to be clearly displayed and does not need to be blurred. In this case, the electronic device can obtain the background mask information but not the mid-ground mask information. If the focus point is in the foreground, it means the mid-ground area needs to be blurred, and the electronic device can obtain both mid-ground and background mask information. Mid-ground and background mask information includes both mid-ground and background mask information.

[0301] In a possible implementation, taking the background image information including background mask information as an example, the electronic device can combine the depth information of the first image and the background depth range, and filter out the region corresponding to the background depth range from the depth information to obtain the background image information. The specific process of obtaining the background image information of the first image can be found in [reference needed]. Figure 5 The relevant description of obtaining background mask information in (2.1) of the corresponding embodiment will not be repeated here.

[0302] S1503. Downsample the information of the foreground image to obtain information of M foreground images with different resolutions, where M is an integer greater than 1.

[0303] In this application embodiment, the downsampling method may include nearest neighbor interpolation, bilinear interpolation, and bicubic interpolation, etc., and this application embodiment does not limit it.

[0304] The specific process of downsampling the foreground image information to obtain information from M foreground images at different resolutions can be found in [reference needed]. Figure 5 The description of the multi-scale processing of foreground information in the corresponding embodiment (1.5) will not be repeated here. M can include integers from 2 to 6.

[0305] S1504. The information of M foreground images with different resolutions is blurred to obtain the information of M blurred foreground images.

[0306] In this embodiment, the process of blurring the information of M foreground images with different resolutions to obtain the information of M blurred foreground images can be referred to... Figure 5 The description of the multi-scale blurring of foreground information in the corresponding embodiment (1.6) will not be repeated here.

[0307] S1505. Perform image fusion on the information of the M blurred foreground images to obtain the fused foreground image.

[0308] In this embodiment of the application, the process of image fusion of information from M blurred foreground images to obtain a fused foreground image can be referred to... Figure 5 The description of the multi-scale fusion of the foreground region in the corresponding embodiment (1.7) will not be repeated here.

[0309] S1506. Blur the background image information to obtain a blurred background image.

[0310] In this embodiment of the application, the blurring process may include blurring using a blurring algorithm, which may include Gaussian blur and bokeh blur. The specific blurring process used is not limited in this embodiment of the application.

[0311] The specific process of blurring the background image to obtain the blurred background image can be found in [reference needed]. Figure 5 The description of the background information blurring in the corresponding embodiment (2.2) will not be repeated here.

[0312] S1507. Perform image fusion on the merged foreground image and the blurred background image to obtain a second image, which is a blurred image.

[0313] In this embodiment of the application, the second image can be understood as an image that has been blurred, and the blurred image includes an image with a blurred foreground.

[0314] The specific process of fusing the merged foreground image and the blurred background image to obtain the second image can be referred to... Figure 5 The corresponding embodiment describes the image fusion of the foreground and mid-ground regions in (3), and Figure 14 The relevant descriptions in the corresponding embodiments will not be repeated here.

[0315] The image processing method provided in this application embodiment can perform blurring processing on the foreground image and the background image respectively, including multi-scale processing and multi-scale blurring of the foreground image, and then merging the foreground image after multi-scale blurring processing and the background image after blurring processing, thereby obtaining an image with blurring effect on both the foreground and the background.

[0316] Optional, in Figure 15 Based on the corresponding embodiment, the fused foreground image and the blurred background image are image fused to obtain a second image, which may include: generating fusion weights for each pixel in M ​​foreground images of different resolutions; and upsampling and image fusion of the M foreground images of different resolutions based on the fusion weights.

[0317] In this application embodiment, the upsampling method may include bilinear interpolation, bicubic interpolation, or convolutional neural network, etc., and this application embodiment does not limit it.

[0318] The specific process of generating the fusion weights for each pixel in M ​​foreground images of different resolutions can be found in [reference needed]. Figure 14 The relevant descriptions of (1) the construction of multi-scale gradients, (2) the calculation of pixel-level fusion weights, (3) the construction of energy functions and (4) the optimization of energy functions and the generation of weights in the corresponding embodiments will not be repeated here.

[0319] Specifically, the process of upsampling and fusing M foreground images of different resolutions based on fusion weights can be found in [reference needed]. Figure 14 The relevant descriptions of (5) single-layer fusion and (6) upsampling and multi-scale fusion in the corresponding embodiments will not be repeated here.

[0320] Understandably, by upsampling and fusing M foreground images of different resolutions based on fusion weights, the image fusion can be made more natural and the edge transitions of the images can be smoother.

[0321] Optional, in Figure 15 Based on the corresponding implementation, generating fusion weights for each pixel in M ​​foreground images of different resolutions may include: calculating pixel-level fusion weights, where the pixel-level fusion weights are the weights of each pixel in the foreground image; and optimizing the pixel-level fusion weights based on a Markov random field (MRF) function to generate fusion weights.

[0322] In this embodiment of the application, the process of calculating pixel-level fusion weights can be referred to Figure 14 The relevant description of the calculation of pixel-level fusion weights in the corresponding embodiment (2) will not be repeated here.

[0323] The process of generating fusion weights by optimizing pixel-level fusion weights based on Markov Random Field (MRF) functions can be found in [reference needed]. Figure 14 The relevant descriptions of (3) the construction of the energy function and (4) the optimization of the energy function and the generation of weights in the corresponding embodiments will not be repeated here.

[0324] Understandably, since pixel-level fusion weights may not have sufficient spatial smoothness, in order to improve the continuity of the foreground fusion region and reduce the fluctuation of fusion weights between pixels, pixel-level fusion weights can be optimized based on Markov Random Field (MRF) functions, thereby obtaining smooth fusion weights at various scales.

[0325] Optional, in Figure 15 Based on the corresponding embodiments, the pixel-level fusion weight is related to the structural information of the foreground image, the hue of the foreground image, and the color of the foreground image.

[0326] In this embodiment, the descriptions of the structural information, hue, and color of the foreground image can be found in the following references: Figure 14 The relevant description of the calculation of pixel-level fusion weights in the corresponding embodiment (2) will not be repeated here.

[0327] It is understandable that pixel-level fusion weights calculated based on the structural information, hue, and color of the foreground image can result in a better display effect in terms of hue and color of the generated image.

[0328] Optional, in Figure 15 Based on the corresponding implementation, the pixel-level fusion weight W_foreground satisfies the following formula:

[0329] W_foreground=W_foreground_structure*W_foreground_tone*W_foreground_color.

[0330] Wherein, W_foreground_structure is the weight corresponding to the structural information of the foreground image, W_foreground_tone is the weight corresponding to the hue of the foreground image, and W_foreground_color is the weight corresponding to the color of the foreground image. W_foreground_structure, W_foreground_tone, and W_foreground_color are all obtained based on the weight generation function Weight_Gen_Func.

[0331] In this embodiment, the calculation method for pixel-level fusion weights can be referred to Figure 14 The relevant description of the calculation of pixel-level fusion weights in the corresponding embodiment (2) will not be repeated here.

[0332] Understandably, the Weight_Gen_Func function can perform Gaussian weighting or smoothing on gradient information, hue statistics, and color statistics, which makes the distribution of the generated pixel-level fusion weights in gradient, hue, and color smoother and more continuous.

[0333] Optional, in Figure 15 Based on the corresponding implementation, the MRF function satisfies the following formula:

[0334] MRF_Fusion_Weight_Pyra=MRF(per_pixel_fusion_weight_Pyra).

[0335] Wherein, MRF_Fusion_Weight_Pyra is the fusion weight of each pixel after optimization by the MRF function, and per_pixel_fusion_weight_Pyra is the weight of each pixel in the M foreground images with different resolutions. per_pixel_fusion_weight_Pyra is calculated based on the pixel-level fusion weight.

[0336] In this embodiment, the specific calculation method of the MRF function can be referred to Figure 14 The relevant description of the calculation of the energy function in (3) of the corresponding embodiment will not be repeated. The MRF function can be used to optimize the weight of each pixel in M ​​foreground images with different resolutions, improve the continuity of the foreground fusion region, reduce the fluctuation of fusion weight between pixels, and thus improve the display effect of the image.

[0337] Optional, in Figure 15 Based on the corresponding embodiment, upsampling and image fusion of M foreground images with different resolutions based on fusion weights can include: using a first fusion method, based on fusion weights, performing image fusion on the foreground image of the Nth layer and the background image of the Nth layer to obtain the image fusion result of the Nth layer, where N is an integer greater than 1 and less than or equal to M; upsampling the image fusion result of the Nth layer to obtain the upsampled result of the Nth layer; performing image fusion on the foreground image of the (N-1)th layer and the background image of the (N-1)th layer to obtain the image fusion result of the (N-1)th layer; and performing image fusion on the upsampled result of the Nth layer and the image fusion result of the (N-1)th layer.

[0338] In this embodiment, the first fusion method may include the Laplace Blending fusion method described in the above embodiments. Image fusion of the upsampling result of layer N and the image fusion result of layer N-1 may include using an image fusion Fusion function to perform image fusion of the upsampling result of layer N and the image fusion result of layer N-1. For example, this Fusion function may perform fusion processing such as weighted averaging, stitching, and convolution.

[0339] For details on the image fusion process using the first fusion method, please refer to [link / reference]. Figure 5 The corresponding embodiment (3.2) describes the fusion of foreground and midground, and Figure 14 The relevant descriptions of upsampling and multi-scale fusion in the corresponding embodiments (6) will not be repeated here.

[0340] Understandably, upsampling an image can generate a higher resolution image. Employing the first fusion method for multi-scale fusion and reconstruction allows for the fusion of different image layers (i.e., images in different frequency bands) using different rules, resulting in images with better foreground blurring.

[0341] Optional, in Figure 15 Based on the corresponding embodiments, the first fusion method satisfies the following formula:

[0342] Fusion_out_layger_N=W_foreground_N*foreground_N+W_background_N*background_N.

[0343] Wherein, Fusion_out_layger_N is the image fusion result of layer N, foreground_N is the foreground image of layer N, W_foreground_N is the fusion weight corresponding to the foreground image of layer N, background_N is the background image of layer N, and W_background_N is the fusion weight corresponding to the background image of layer N.

[0344] In this embodiment, the calculation method for the first fusion method can be referred to... Figure 14 The descriptions of upsampling and multi-scale fusion in the corresponding embodiments (6) will not be repeated here. It is understandable that by defining the weights of the low-resolution layer and the current layer through weights, more use of the low-resolution foreground can achieve a more blurred foreground effect.

[0345] Optional, in Figure 15Based on the corresponding embodiment, the fused foreground image and the blurred background image are image-fused to obtain a second image. This can include: using a second fusion method to perform image fusion of the fused foreground image and the blurred background image to obtain a second image; the second fusion method satisfies the following formula:

[0346] Fusion_Result=Front_Scene_Bokeh*Front_Scene_Fusion_Mask+Back_Scene_Bokeh*(1-Front_Scene_Fusion_Mask).

[0347] Wherein, Fusion_Result is the second image, Front_Scene_Bokeh is the fused foreground image, Back_Scene_Bokeh is the blurred background image, and Front_Scene_Fusion_Mask is the fusion mask, which is used to represent the weight ratio of the fused foreground image in the second image.

[0348] In this embodiment, the second fusion method may include the Alpha Blending fusion method described in the above embodiments. The specific process of image fusion using the second fusion method can be found in [reference needed]. Figure 5 The relevant description of the fusion of foreground and midground in the corresponding embodiment (3.2) will not be repeated here.

[0349] Understandably, the second fusion method can control the fusion ratio of two images by using transparency. The formula for the second fusion method is relatively simple, easy to implement, and has relatively low computational complexity. When performing image fusion, it is more convenient to achieve a smooth transition effect.

[0350] Optional, in Figure 15 Based on the corresponding embodiments, the first fusion method includes Laplace Blending fusion method, and the second fusion method includes Alpha Blending fusion method.

[0351] Understandably, the first fusion method achieves image fusion through multi-scale decomposition and reconstruction, allowing for smooth edge transitions in the fused image and preserving more image details. The second fusion method controls the fusion ratio of the two images using transparency (alpha value), thus achieving a smooth transition effect during image fusion.

[0352] Optional, in Figure 15Based on the corresponding embodiments, obtaining information about the foreground image and the background image of the first image may include: obtaining target information based on the depth map and the focus position of the first image, wherein the target information includes information about the foreground image and information about the background image; wherein the target information is used to distinguish between the foreground image and the background image in the first image, the foreground image being the image before the focus position in the first image, and the background image being the image after the focus position in the first image.

[0353] In this embodiment of the application, the target information may include mask information, and the information of the foreground image and the background image can be referred to the relevant description in step S1502, which will not be repeated here.

[0354] The process of acquiring information from the foreground image can be referred to Figure 5 The descriptions of (1.1) the calculation of dual-camera depth, (1.2) the acquisition of the focus position, (1.3) the calculation of the depth of field range using simulated aperture settings, and (1.4) the acquisition of foreground mask information in the corresponding embodiments will not be repeated here. The process of acquiring background image information can be found in [reference needed]. Figure 5 The relevant description of obtaining background mask information in (2.1) of the corresponding embodiment will not be repeated here.

[0355] Based on information from the foreground image, electronic devices can perform multi-scale processing to obtain foreground information at different resolutions. Furthermore, they can blur information from both the foreground and background images, resulting in images where both the foreground and background have a blurred effect, thus enhancing the overall image blurring display.

[0356] Optional, in Figure 15 Based on the corresponding embodiments, the shooting operation includes shooting operations in portrait mode scenes.

[0357] Understandably, blurring the foreground and background images in portrait mode can more effectively highlight the subject, making it appear clearer and more three-dimensional in the image, and enhancing the overall sense of space.

[0358] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties. Furthermore, the collection, use and processing of the relevant data must comply with the relevant laws, regulations and standards of the relevant countries and regions, and corresponding operation portals are provided for users to choose to authorize or refuse.

[0359] The foregoing primarily describes the solutions provided by the embodiments of this application from a methodological perspective. To achieve the aforementioned functions, it includes corresponding hardware structures and / or software modules for executing each function. Those skilled in the art should readily recognize that, based on the method steps of the examples described in conjunction with the embodiments disclosed herein, this application can be implemented in hardware or a combination of hardware and computer software. Whether a function is executed in hardware or by computer software driving hardware depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0360] This application embodiment can divide the apparatus for implementing the method into functional modules based on the above method examples. For example, each function can be divided into its own functional modules, or two or more functions can be integrated into one processing module. The integrated module can be implemented in hardware or as a software functional module. It should be noted that the module division in this application embodiment is illustrative and only represents one logical functional division; other division methods may be used in actual implementation.

[0361] like Figure 16 The diagram shows a chip structure according to an embodiment of this application. The chip 1600 includes one or more processors 1601, communication lines 1602, communication interfaces 1603, and memory 1604.

[0362] In some implementations, memory 1604 stores elements such as executable modules or data structures, or subsets thereof, or extended sets thereof.

[0363] The methods described in the embodiments of this application can be applied to, or implemented by, processor 1601. Processor 1601 may be an integrated circuit chip with signal processing capabilities. During implementation, each step of the above methods can be completed by integrated logic circuits in the hardware of processor 1601 or by instructions in software form. Processor 1601 may be a general-purpose processor (e.g., a microprocessor or conventional processor), a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gates, transistor logic devices, or discrete hardware components. Processor 1601 can implement or execute the various processing-related methods, steps, and logic block diagrams disclosed in the embodiments of this application.

[0364] The steps of the method disclosed in the embodiments of this application can be directly implemented by a hardware decoding processor, or implemented by a combination of hardware and software modules in the decoding processor. The software modules can be located in mature storage media in the art, such as random access memory, read-only memory, programmable read-only memory, or electrically erasable programmable read-only memory (EEPROM). This storage medium is located in memory 1604, and processor 1601 reads information from memory 1604 and, in conjunction with its hardware, completes the steps of the above method.

[0365] The processor 1601, memory 1604 and communication interface 1603 can communicate with each other via communication line 1602.

[0366] In the above embodiments, the instructions stored in the memory for execution by the processor can be implemented in the form of a computer program product. This computer program product can be pre-written into the memory, or it can be downloaded and installed into the memory as software.

[0367] This application also provides a computer program product comprising one or more computer instructions. When the computer program instructions are loaded and executed on a computer, all or part of the flow or function according to the embodiments of this application is generated. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions may be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, computer instructions may be transmitted from a website site, computer, server, or data center to another website site, computer, server, or data center via wired (e.g., coaxial cable, fiber optic, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium may be any available medium that a computer can store or a data storage device such as a server or data center that integrates one or more available media. For example, available media may include magnetic media (e.g., floppy disk, hard disk, or magnetic tape), optical media (e.g., digital versatile disc (DVD)), or semiconductor media (e.g., solid-state disk (SSD)).

[0368] This application also provides a computer-readable storage medium. The methods described in the above embodiments can be implemented, in whole or in part, by software, hardware, firmware, or any combination thereof. The computer-readable medium may include computer storage media and communication media, and may also include any medium capable of transferring a computer program from one place to another. The storage medium can be any target medium accessible by a computer.

[0369] As one possible design, computer-readable media may include compact disc read-only memory (CD-ROM), RAM, ROM, EEPROM, or other optical disc storage; computer-readable media may also include disk storage or other disk storage devices. Furthermore, any connecting cable may also be appropriately referred to as computer-readable media. For example, if software is transmitted from a website, server, or other remote source using coaxial cable, fiber optic cable, twisted pair, DSL, or wireless technologies such as infrared, radio, and microwave, then coaxial cable, fiber optic cable, twisted pair, DSL, or wireless technologies such as infrared, radio, and microwave are included in the definition of media. As used herein, disks and optical discs include optical discs (CD), laser discs, optical discs, digital versatile discs (DVD), floppy disks, and Blu-ray discs, where disks typically reproduce data magnetically, while optical discs optically reproduce data using lasers.

[0370] This application describes embodiments with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this application. It should be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processing unit of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processing unit of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart illustrations. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.

Claims

1. An image processing method, characterized in that, The method includes: In response to the shooting operation, acquire the first image; Obtain information about the foreground image and the background image of the first image; The information of the foreground image is downsampled to obtain information of M foreground images with different resolutions, where M is an integer greater than 1; The information of the M foreground images with different resolutions is blurred to obtain the information of the M blurred foreground images; Image fusion is performed on the information of the M blurred foreground images to obtain the fused foreground image; The background image information is blurred to obtain a blurred background image; The fused foreground image and the blurred background image are merged to obtain a second image, which is a blurred image. The step of fusing the fused foreground image and the blurred background image to obtain a second image includes: Generate the fusion weights of each pixel in the M fused foreground images with different resolutions; Based on the fusion weights, the M fused foreground images with different resolutions and the blurred background images are upsampled and fused to obtain the second image.

2. The method according to claim 1, characterized in that, The fusion weights for each pixel in the M fused foreground images of different resolutions are generated, including: Calculate pixel-level fusion weights, where the pixel-level fusion weights are the weights of each pixel in the fused foreground image; The pixel-level fusion weights are generated by optimizing the Markov Random Field (MRF) function.

3. The method according to claim 2, characterized in that, The pixel-level fusion weights are related to the structural information of the foreground image, the hue of the foreground image, and the color of the foreground image.

4. The method according to claim 2, characterized in that, The pixel-level fusion weight W_foreground satisfies the following formula: W_foreground = W_foreground_structure W_foreground_tone W_foreground_color; Wherein, W_foreground_structure is the weight corresponding to the structural information of the fused foreground image, W_foreground_tone is the weight corresponding to the hue of the fused foreground image, and W_foreground_color is the weight corresponding to the color of the fused foreground image. W_foreground_structure, W_foreground_tone, and W_foreground_color are all obtained based on the weight generation function Weight_Gen_Func.

5. The method according to claim 2, characterized in that, The MRF function satisfies the following formula: MRF_Fusion_Weight_Pyra = MRF(per_pixel_fusion_weight_Pyra); Wherein, MRF_Fusion_Weight_Pyra is the fusion weight of each pixel after optimization by the MRF function, and per_pixel_fusion_weight_Pyra is the weight of each pixel in the fused foreground image of the M different resolutions, and per_pixel_fusion_weight_Pyra is calculated based on the pixel-level fusion weight.

6. The method according to any one of claims 1-3 and 5, characterized in that, The step of upsampling and image fusion processing on the M fused foreground images of different resolutions based on the fusion weights to obtain a second image includes: Using the first fusion method, based on the fusion weight, the fused foreground image of the Nth layer and the blurred background image of the Nth layer are fused to obtain the image fusion result of the Nth layer, where N is an integer greater than 1 and less than or equal to M; The image fusion result of the Nth layer is upsampled to obtain the upsampled result of the Nth layer; Image fusion is performed on the fused foreground image of layer N-1 and the blurred background image of layer N-1 to obtain the image fusion result of layer N-1. The upsampling result of the Nth layer and the image fusion result of the (N-1)th layer are subjected to image fusion processing to obtain the second image.

7. The method according to claim 6, characterized in that, The first fusion method satisfies the following formula: Fusion_out_layger_N = W_foreground_N foreground_N + W_background_N background_N; Wherein, Fusion_out_layger_N is the image fusion result of the Nth layer, foreground_N is the fused foreground image of the Nth layer, W_foreground_N is the fusion weight corresponding to the fused foreground image of the Nth layer, background_N is the blurred background image of the Nth layer, and W_background_N is the fusion weight corresponding to the virtualized background image of the Nth layer.

8. The method according to claim 1, characterized in that, The step of fusing the fused foreground image and the blurred background image to obtain a second image includes: The second fusion method is used to fuse the fused foreground image and the blurred background image to obtain the second image. The second fusion method satisfies the following formula: Fusion_Result = Front_Scene_Bokeh Front_Scene_Fusion_Mask + Back_Scene_Bokeh (1 - Front_Scene_Fusion_Mask); Wherein, Fusion_Result is the second image, Front_Scene_Bokeh is the fused foreground image, Back_Scene_Bokeh is the blurred background image, and Front_Scene_Fusion_Mask is the fusion mask, which is used to represent the weight ratio of the fused foreground image in the second image.

9. The method according to claim 6, characterized in that, The first fusion method includes the Laplace Blending fusion method.

10. The method according to claim 8, characterized in that, The second fusion method includes the AlphaBlending fusion method.

11. The method according to any one of claims 1-3, 5, and 7-10, characterized in that, The step of obtaining information about the foreground image and the background image of the first image includes: Based on the depth map and focus position of the first image, target information is obtained, including information of the foreground image and information of the background image; The target information is used to distinguish between the foreground image and the background image in the first image. The foreground image is the image before the focus position in the first image, and the background image is the image after the focus position in the first image.

12. The method according to any one of claims 1-3, 5, and 7-10, characterized in that, The shooting operation includes shooting operations in portrait mode scenes.

13. An electronic device, characterized in that, The electronic device includes: one or more processors and memory; The memory is coupled to the one or more processors, the memory being used to store computer program code, the computer program code including computer instructions, the one or more processors invoking the computer instructions to cause the electronic device to perform the method as described in any one of claims 1-12.

14. A chip system, characterized in that, The chip system is applied to an electronic device, the chip system including one or more processors, the one or more processors being configured to invoke computer instructions to cause the electronic device to perform the method as described in any one of claims 1-12.

15. A computer-readable storage medium, characterized in that, The computer-readable storage medium includes computer instructions that, when executed on an electronic device, cause the electronic device to perform the method as described in any one of claims 1-12.

16. A computer program product, characterized in that, The computer program product includes computer program code that, when run on an electronic device, causes the electronic device to perform the method as described in any one of claims 1-12.

Citation Information

Patent Citations

  • Background blurring method and device, terminal equipment and computer readable storage medium

    CN111127303A

  • Picture background blurring method and device, computer equipment and storage medium

    CN113129207A

  • Image bokeh method based on self-supervised multi-scale pyramid fusion network

    CN114757860A