Image processing method and device

By collecting and processing images in electronic devices, determining the location and transition areas of edge features, and fusing these areas into the blurred image, the problem of edge error or background error in the blurred image is solved, and the blurred effect is improved.

CN120070156AActive Publication Date: 2025-05-30HONOR DEVICE CO LTD
View PDF 15 Cites 0 Cited by

Patent Information

Application Number
CN202311588117.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-11-23
Publication Date
2025-05-30
Estimated Expiration
2043-11-23

AI Technical Summary

Technical Problem

Due to the limitations of the hardware module of electronic devices, blurred images may cause false edges or backgrounds to be incorrectly clear.

Method used

By collecting the first image, it obtains its corresponding blur image, the position of edge features, and the transition area composed of edge features. Then, a first area composed of edge features is determined from the first image using the position of edge features, and the area is fused into the blurred image to obtain a target image.

Benefits of technology

It solves the situation where edges are false or backgrounds are wrongly clear in the blurred image, and improves the blurring effect of edge features.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120070156A_ABST
    Figure CN120070156A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides an image processing method and device, and relates to the technical field of terminals, so that an electronic device can obtain a mask image corresponding to a clear image and a ternary image corresponding to the clear image, determine the position of an edge feature through the mask image, determine a transition region comprising the edge feature through the ternary image, and determine the position of the edge feature through the transition region. Therefore, the electronic device can obtain the edge region corresponding to the position of the edge feature from the clear image based on the position of the edge feature, and fuse the edge region into the blurred image, thereby solving the problem that the edge is false or the background is false clear in the blurred image.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of terminal technologies, and in particular, to an image processing method and apparatus. Background Art

[0002] With the popularization and development of the Internet, people's functional requirements for electronic devices have become increasingly diverse. For example, an electronic device can not only support a shooting function, but also support blurring processing of the captured image, enabling the user to see a blurred image with a clear foreground and a blurred background, giving the captured image a better sense of space.

[0003] However, due to the limitations of the electronic device's hardware module, the blurred image may have the situation of edge mis-blurring or background mis-clarity. Summary of the Invention

[0004] Embodiments of this application provide an image processing method and apparatus for enhancing the blurring effect of edge features.

[0005] In a first aspect, embodiments of this application provide an image processing method, the method includes: in response to a photographing operation, acquiring a first image; obtaining a blurred image corresponding to the first image, positions of edge features in the first image, and a transition region formed by the edge features in the first image; wherein, the blurred image is obtained by blurring the background in the first image; using the positions of the edge features to determine a first region formed by the edge features from the first image, and fusing the first region into the blurred image to obtain a target image; wherein, the first region is within the transition region.

[0006] An electronic device can accurately select a first region formed by clear edge features from the first image based on the positions of the edge features and the first region, and solve the situation of edge mis-blurring or background mis-clarity in the blurred image by fusing the first region into the blurred image.

[0007] In a possible implementation, before obtaining the positions of edge features in the first image and the transition region formed by the edge features in the first image, the method further includes: obtaining a mask image corresponding to the first image and a ternary image corresponding to the first image, the mask image corresponding to the first image includes the positions of the edge features, and the ternary image corresponding to the first image includes the transition region.

[0008] The mask image corresponding to the first image may be the first mask image described in the embodiments of this application, and the ternary image corresponding to the first image may be the first ternary image described in the embodiments of this application.

[0009] Enable the electronic device to obtain the mask image corresponding to the clear image and the ternary image corresponding to the clear image, determine the position of the edge feature through the mask image, determine the transition region including the edge feature through the ternary image, and then the electronic device can obtain the edge region corresponding to the position of the edge feature from the clear image based on the position of the edge feature, and fuse the edge region into the blurred image to solve the situation of incorrect blurring of the edge or incorrect clarity of the background in the blurred image.

[0010] In a possible implementation manner, obtaining the mask image corresponding to the first image and the ternary image corresponding to the first image includes: inputting the first image into a first model, and outputting the mask image corresponding to the first image and the ternary image corresponding to the first image; wherein, the first model includes: a shared editor, a semantic segmentation branch, a detail prediction branch, and a fusion branch, the shared editor is used to extract the image features of the first image, the semantic segmentation branch is used to perform semantic segmentation on the first image and obtain the foreground in the first image, the detail prediction branch is used to obtain the edge features in the first image, and the fusion branch is used to fuse the foreground in the first image and the edge features in the first image to obtain the mask image corresponding to the first image and the ternary image corresponding to the first image.

[0011] The electronic device can output a mask image that can identify clear edges and a ternary image that can frame the edge features within the transition region through the target neural network model, so that the accurate position and range of the edge features can be determined using the mask image and the ternary image.

[0012] In a possible implementation manner, the method further includes: the first model is trained in the following manner: inputting the training image into the shared editor to obtain the image features of the training image; inputting the image features of the training image into the semantic segmentation branch to obtain the output result of the semantic segmentation branch; obtaining the Laplacian high-frequency map corresponding to the training image, inputting the Laplacian high-frequency map into the detail prediction branch to obtain the output result of the detail prediction branch; inputting the output result of the semantic segmentation branch and the output result of the detail prediction branch into the fusion branch to obtain the predicted mask image and the predicted ternary image; when the difference between the predicted mask image and the mask image ground truth satisfies the first loss function, and the difference between the predicted ternary image and the ternary image ground truth satisfies the second loss function, the trained first model is obtained.

[0013] The training device can output the foreground contour of the image that can be recognized through the semantic segmentation branch, determine the range of the edge features based on the detail prediction branch, and obtain a clear edge contour and a region including the edge features through the fusion branch.

[0014] In a possible implementation, the method further includes: performing Gaussian blurring on the edges of the ground truth mask map to obtain a second image; obtaining the trained semantic segmentation branch when the difference between the second image and the output result of the semantic segmentation branch satisfies a fourth loss function.

[0015] The second image may be the mask edge blurred map described in the embodiments of the present application, and the fourth loss function may be loss4 described in the embodiments of the present application.

[0016] Since the semantic segmentation branch is relatively rough in semantic segmentation of images and can only obtain a general portrait foreground, while the edges of the ground truth mask map are relatively accurate, in order to enable the semantic segmentation branch to output an image consistent with the foreground contour in the ground truth mask map, the training device may perform loss4 calculation on the ground truth mask map and the output result of the semantic segmentation branch.

[0017] In a possible implementation, the method further includes: obtaining the image within the transition region from the ground truth mask map according to the transition region in the ternary ground truth map to obtain a third image; obtaining the trained detail prediction branch when the difference between the third image and the output result of the detail prediction branch satisfies a third loss function.

[0018] The third image is the transition region mask map described in the embodiments of the present application, and the third loss function may be loss3 described in the embodiments of the present application.

[0019] In order to enable the detail prediction branch to predict more details within the transition region containing edge features, the training device may perform loss3 calculation on the ternary ground truth map and the output result of the detail prediction branch.

[0020] In a possible implementation, the method further includes: obtaining a second region containing edge features from the training image based on a semantic segmentation method; performing connected component screening on the second region to obtain a third region; performing dilation on the third region to obtain a fourth region; performing dilation on the ground truth mask map to obtain a dilated mask map; obtaining the intersection of the dilated mask map and the fourth region to obtain a first transition region; determining the ternary ground truth map based on the first transition region and the foreground in the dilated mask map.

[0021] The electronic device can determine the position of the edge features through connected component screening and determine the region containing edge features through operations such as dilation to obtain a relatively accurate ternary ground truth map.

[0022] Moreover, compared with directly dilating the mask image conventionally to obtain a ternary image in which the transition region wraps the entire image, the transition region of the ternary image truth value in this application only contains edge feature distributions.

[0023] In a possible implementation manner, the edge features include one or more of the following: the hair of a person, the wool in a sweater, or the hair of an animal, or the branches and leaves of a plant.

[0024] In a possible implementation manner, the collecting the first image in response to the photographing operation includes: in response to the photographing operation, obtaining at least two images with different exposure times, and performing image fusion on the at least two images with different exposure times to obtain the first image.

[0025] The electronic device can obtain the first image with HDR effect through image fusion between at least two images with different exposure times.

[0026] In a second aspect, an embodiment of the present application provides an image processing method, and the method includes: the first model is trained in the following manner: inputting a training image into the shared editor to obtain the image features of the training image; inputting the image features of the training image into the semantic segmentation branch to obtain the output result of the semantic segmentation branch; obtaining the Laplacian high-frequency map corresponding to the training image, and inputting the Laplacian high-frequency map into the detail prediction branch to obtain the output result of the detail prediction branch; inputting the output result of the semantic segmentation branch and the output result of the detail prediction branch into the fusion branch to obtain a predicted mask image and a predicted ternary image; when the difference between the predicted mask image and the mask image truth value satisfies the first loss function, and the difference between the predicted ternary image and the ternary image truth value satisfies the second loss function, the trained first model is obtained.

[0027] Among them, the main body for training the first model can be an electronic device or a server, and this application embodiment does not make a limitation on this.

[0028] In a possible implementation manner, the method further includes: performing Gaussian blur processing on the edge of the mask image truth value to obtain a second image; when the difference between the second image and the output result of the semantic segmentation branch satisfies the fourth loss function, the trained semantic segmentation branch is obtained.

[0029] In a possible implementation manner, the method further includes: obtaining the image in the transition region from the mask image truth value according to the transition region in the ternary image truth value to obtain a third image; when the difference between the third image and the output result of the detail prediction branch satisfies the third loss function, the trained detail prediction branch is obtained.

[0030] In a possible implementation, the method further includes: obtaining a second region containing edge features from the training image based on a semantic segmentation method; performing connected component filtering on the second region to obtain a third region; performing dilation on the third region to obtain a fourth region; performing dilation on the ground truth mask image to obtain a dilated mask image; obtaining an intersection of the dilated mask image and the fourth region to obtain a first transition region; and determining the ground truth ternary image based on the first transition region and the foreground in the dilated mask image.

[0031] In a possible implementation, the edge features include one or more of the following: hair of a human portrait, wool in a sweater, or hair of an animal, or branches and leaves of a plant.

[0032] In a third aspect, an embodiment of the present application provides an image processing device, which includes a display unit and a processing unit. The display unit is configured to process the steps of data display in the image processing device, and the processing unit is configured to process the steps of data processing in the image processing device.

[0033] In a possible implementation, the image processing device may further include a storage unit. The storage unit may include one or more memories, and the memory may be a device or a circuit in one or more devices for storing programs or data.

[0034] In a fourth aspect, an embodiment of the present application provides an electronic device, which includes a processor and a memory. The memory is used to store code instructions; the processor is used to run the code instructions so that the electronic device executes the method described in the first aspect or any implementation manner of the first aspect.

[0035] In a fifth aspect, an embodiment of the present application provides a computer-readable storage medium, which stores instructions. When the instructions are executed, the computer executes the method described in the first aspect or any implementation manner of the first aspect.

[0036] In a sixth aspect, a computer program product includes a computer program. When the computer program is run, the computer executes the method described in the first aspect or any implementation manner of the first aspect.

[0037] In a seventh aspect, a chip includes: a processor, configured to read instructions stored in a storage memory. When the processor executes the instructions, the chip executes the method described in the first aspect or any implementation manner of the first aspect.

[0038] It should be understood that the technical solutions of the second to seventh aspects of this application correspond to those of the first aspect of this application, and the beneficial effects achieved by each aspect and the corresponding feasible implementation manners are similar, and will not be elaborated herein. BRIEF DESCRIPTION OF THE DRAWINGS

[0039] Figure 1 FIG. 1 is a schematic diagram of a scenario provided by an embodiment of this application;

[0040] Figure 2 FIG. 2 is a schematic diagram of the hardware structure of an electronic device provided by an embodiment of this application;

[0041] Figure 3 FIG. 3 is a schematic diagram of the software structure of the electronic device provided by an embodiment of this application;

[0042] Figure 4 FIG. 4 is a schematic flowchart of an image processing method provided by an embodiment of this application;

[0043] Figure 5 FIG. 5 is a schematic diagram of the architecture of a neural network model provided by an embodiment of this application;

[0044] Figure 6 FIG. 6 is a schematic diagram of an image provided by an embodiment of this application;

[0045] Figure 7 FIG. 7 is a schematic diagram of the steps for generating a ternary graph provided by an embodiment of this application;

[0046] Figure 8 FIG. 8 is a schematic flowchart of another image processing method provided by an embodiment of this application;

[0047] Figure 9 FIG. 9 is a schematic diagram of the structure of an image processing apparatus provided by an embodiment of this application;

[0048] Figure 10 FIG. 10 is a schematic diagram of the hardware structure of another electronic device provided by an embodiment of this application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0049] For the convenience of clearly describing the technical solutions of the embodiments of this application, the nouns involved in the embodiments of this application are explained. It can be understood that this explanation is for more clearly explaining the embodiments of this application and does not necessarily constitute a limitation on the embodiments of this application.

[0050] 1. Preview Image and Captured Image

[0051] The preview image can be data that is collected in real time by the camera of the electronic device and is allowed to be displayed in the preview screen. For example, when the electronic device receives an operation from the user to open the camera application, the electronic device can collect the preview image captured by the camera and display it in real time in the preview screen of the camera application.

[0052] The captured image can be data obtained based on the capture button in the electronic device, such as the target image described in the embodiments of the present application. For example, when the electronic device receives a trigger operation from the user for the capture button, the electronic device can obtain the captured image acquired by the camera at the moment of capture.

[0053] 2. High Dynamic Range (HDR)

[0054] HDR is a processing technology that improves the brightness and contrast of images. Compared with ordinary images, HDR can provide more dynamic range and image details. By synthesizing the final HDR image using the images with the best details corresponding to each exposure time, it can better reflect the visual effect in the real environment.

[0055] A possible implementation for the electronic device to determine whether the current is an HDR scene is as follows: The electronic device downsamples the preview image by 4 times to obtain a preview thumbnail, and determines whether the proportion of the number of highlight pixels in the preview thumbnail to the total number of all pixels in the preview thumbnail is greater than a preset pixel threshold. Among them, the preview thumbnail can be obtained by retaining one row of pixel points out of every two rows of pixel points in the picture corresponding to the preview image, and storing every other column of the obtained pixel points in each row. The highlight pixel can be determined based on the gray threshold of the pixel point, and the pixel threshold can be used to determine whether the current scene is a high-dynamic scene. Among them, the electronic device can determine whether the current is a high-dynamic scene based on one frame of data or multiple frames of data, and this is not limited in the embodiments of the present application.

[0056] 3. Camera Depth of Field

[0057] The depth of field can be understood as the imaging range in a camera lens or other imager that can obtain a clear image, or can be understood as the clear range formed before and after the focus point. Among them, the focus point can include the clearest point obtained when light passes through the lens and focuses on the photosensitive element. The front depth of field can include the clear range before the focus point, and the background (or called the rear depth) can include the clear range after the focus point.

[0058] The important factors affecting the depth of field can include the aperture, the lens, and the distance from the object being photographed, etc. When the aperture is larger (the aperture value F is smaller), the depth of field is shallower; when the aperture is smaller (the aperture value F is larger), the depth of field is deeper. When the lens focal length is longer, the depth of field is shallower; when the lens focal length is shorter, the depth of field is deeper.

[0059] 4. Exposure Time (or Exposure Duration)

[0060] The exposure time is the time for which the shutter needs to be opened to project light onto the photosensitive surface of the photographic light-sensitive material, or it can also be understood as the time interval from when the shutter opens to when it closes.

[0061] The exposure time refers to the photosensitive time of the negative. The longer the exposure time, the brighter the photo generated on the negative. Conversely, the darker it is. In the case of relatively dim external light, it is generally required to extend the exposure time to obtain a brighter image.

[0062] 5. Alpha mask image (or simply referred to as mask image)

[0063] The mask image can be an image generated by performing occlusion processing on an image (all or part of it). The mask image can be used to extract the region of interest. Taking the data format of the mask image as 8bit as an example, the alpha value of each pixel point in the mask image is illustrated. In the mask image, the alpha value of the pixel points is distributed between 0 and 255. The foreground region of interest can be set to alpha = 255, and the background region can be set to alpha = 0.

[0064] Usually, the mask image can be used to solve the problem of image matting. For example, the problem of image matting can be modeled as:

[0065] I = alpha × Fg + (255 - alpha) × Bg

[0066] Where, I is the complete image, Fg is the portrait foreground, Bg is the background, and alpha is the mask of the portrait foreground Fg. The value of alpha is distributed between 0 and 255. The region with alpha value of 0 can be the background region, the region with alpha value of 255 can be the foreground region, and the region with alpha value greater than 0 and less than 255 is the transition region in the mask image.

[0067] 6. Trimap

[0068] The trimap can be an image containing three markings: foreground, background, and the mixed region of foreground and background. Each pixel point region in the trimap can be one of 0, 128, and 255. The trimap can achieve a rough division of a given image. For example, the region where trimap = 0 can be the background region, the region where trimap = 255 can be the foreground region, and the region where trimap = 128 can be the transition region.

[0069] 7. Electronic device

[0070] An electronic device can also be referred to as a terminal, (terminal), user equipment (UE), mobile station (MS), mobile terminal (MT), etc. The electronic device can be a mobile phone with a touch screen, smart TV, wearable device, tablet computer (Pad), computer with wireless transceiver function, virtual reality (VR) electronic device, augmented reality (AR) electronic device, wireless terminal in industrial control, wireless terminal in self-driving, wireless terminal in remote medical surgery, wireless terminal in smart grid, wireless terminal in transportation safety, wireless terminal in smart city, wireless terminal in smart home, and so on. The embodiments of the present application do not limit the specific technologies and specific device forms adopted by the electronic device.

[0071] 8. Others

[0072] For the convenience of clearly describing the technical solutions of the embodiments of the present application, in the embodiments of the present application, terms such as "first" and "second" are used to distinguish identical or similar items with basically the same functions and roles. For example, the first value and the second value are only used to distinguish different values, and do not limit their order. Those skilled in the art can understand that terms such as "first" and "second" do not limit the quantity and execution order, and the terms "first" and "second" do not necessarily limit being different.

[0073] It should be noted that in the present application, words such as "exemplary" or "for example" are used to represent examples, illustrations or explanations. Any embodiment or design solution described as "exemplary" or "for example" in the present application should not be construed as being more preferred or more advantageous than other embodiments or design solutions. Exactly speaking, using words such as "exemplary" or "for example" aims to present relevant concepts in a specific manner.

[0074] In this application, "at least one" means one or more, and "a plurality" means two or more. "And / or" describes the association relationship of associated objects, indicating that there can be three relationships. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone, where A and B can be singular or plural. The character " / " generally represents an "or" relationship between the associated objects before and after. "At least one of the following" or its similar expression refers to any combination of these items, including any combination of single item or plural items. For example, at least one of a, b, or c can represent: a, b, c, a and b, a and c, b and c, or a, b, and c, where a, b, and c can be single or multiple.

[0075] In daily photography, keeping the foreground in focus and blurring the background is a common photography method. For the camera function of a mobile phone, due to hardware module limitations, it is difficult for the directly captured image to achieve a sufficient degree of background blurring. An algorithm is required to extract the foreground from the entire image and then blur the background. Therefore, the accuracy of matte extraction for edge features greatly affects the effect of the actually blurred image. However, due to the complex distribution and rich details of edge features, the blurred image is very likely to have cases of incorrect blurring of the edges or incorrect sharpness of the background.

[0076] Exemplarily, Figure 1 is a schematic diagram of a scenario provided for an embodiment of this application. In Figure 1 the corresponding embodiment, taking the electronic device as a mobile phone and the edge feature as hair as examples for illustration, this example does not limit the embodiments of this application.

[0077] As Figure 1 shown in scenario a, this scenario may include a person in the foreground, as well as the sun and sun rays in the background, etc. At least one hair region may be displayed around the person's head, such as hair region 101, and hair region 102, etc.

[0078] In response to the user opening the portrait photography function in the camera, the electronic device may display an interface as Figure 1 shown in b, and this interface may be a preview interface in the portrait photography function. This interface may include one or more of the following: a capture button 103, a preview screen, a control for switching the camera, or a control for indicating the use of the photography function, etc. The camera application may also have other functions in addition to the portrait photography function, such as supporting multiple functions such as aperture photography, night scene photography, video recording, short video, and HDR photography function.

[0079] Among them, the preview image 104 displayed in the preview screen can be obtained by the electronic device identifying the foreground and background of the captured screen and blurring the identified background. In possible implementation manners, the preview image after blurring processing may not be displayed in the preview screen, and this is not limited in the embodiments of the present application.

[0080] In response to a triggering operation of the user on the capture button 103, the electronic device captures the scene shown in a of Figure 1 and obtains the captured image 105 in the interface shown in c of Figure 1 . Among them, the captured image 105 can be obtained by the electronic device identifying the foreground and background of the captured screen and blurring the identified background.

[0081] The captured image 105 may include: a clear hair region 101', a blurred hair region 102', and a clear partial background region 106, etc. Refer to the scene in a of Figure 1 and the captured image shown in c of Figure 1 . The hair region 101 can be a part of the foreground region, so the hair region 101' can be a clear region; the hair region 102 can be a part of the foreground. Since the electronic device misidentifies the foreground, the hair region 102 is misidentified as the background, so that the hair region 102' is blurred; the sun and its rays can be the background. Since the electronic device misidentifies the background, some of the sun rays are misidentified as the foreground, so that some of the background region 106 is not blurred.

[0082] It can be understood that due to the limited ability of the electronic device to perform matte extraction for edge features, the blurred image is very likely to cause incorrect blurring of the edge or incorrect clarity of the background, affecting the clarity of the captured image.

[0083] In view of this, the embodiments of the present application provide an image processing method, enabling the electronic device to obtain a mask image corresponding to a clear image and a ternary image corresponding to the clear image, determining the position of the edge feature through the mask image, determining the transition region including the edge feature through the ternary image, and further enabling the electronic device to obtain an edge region corresponding to the position of the edge feature from the clear image based on the position of the edge feature and fuse the edge region into the blurred image. Among them, the edge region is located within the transition region, and the blurred image can be obtained by blurring the background of the clear image.

[0084] Therefore, in order to better understand the embodiments of the present application, the structure of the electronic device in the embodiments of the present application will be introduced below. Exemplarily, Figure 2 is a schematic diagram of the hardware structure of an electronic device provided by the embodiments of the present application.

[0085] The electronic device may include a processor 110, an external memory interface 120, an internal memory 121, a universal serial bus (USB) interface 130, a charging management module 140, a power management module 141, antenna 1, antenna 2, a mobile communication module 150, a wireless communication module 160, an audio module 170, a speaker 170A, a receiver 170B, a microphone 170C, a headphone jack 170D, a sensor module 180, a button 190, an indicator 192, a camera 193, and a display screen 194, etc.

[0086] Among them, the sensor module 180 may include one or more of the following: a pressure sensor, a gyroscope sensor, a barometric pressure sensor, a magnetic sensor, an acceleration sensor, a distance sensor, a proximity light sensor, a fingerprint sensor, a temperature sensor, a touch sensor, an ambient light sensor, or a bone conduction sensor ( Figure 2 not shown in

[0087] It can be understood that the structure illustrated in the embodiments of the present application does not constitute a specific limitation on the electronic device. In other embodiments of the present application, the electronic device may include more or fewer components than those shown, or combine certain components, or split certain components, or have different component arrangements. The illustrated components may be implemented in hardware, software, or a combination of software and hardware.

[0088] The processor 110 may include one or more processing units. Among them, different processing units may be independent devices or integrated in one or more processors. A memory may also be provided in the processor 110 for storing instructions and data. For example, the processor 110 is used to implement the steps executed in the image processing method provided in the embodiments of the present application, and store the instructions and data related to the image processing method.

[0089] The charging management module 140 is used to receive a charging input from a charger. Among them, the charger may be a wireless charger or a wired charger. The power management module 141 is used to connect the charging management module 140 and the processor 110.

[0090] The wireless communication function of the electronic device may be implemented by antenna 1, antenna 2, the mobile communication module 150, the wireless communication module 160, a modulation and demodulation processor, and a baseband processor, etc.

[0091] The electronic device realizes the display function through a GPU, the display screen 194, and an application processor, etc. The GPU is a microprocessor for image processing, connected to the display screen 194 and the application processor. The GPU is used to perform mathematical and geometric calculations for graphics rendering. For example, the GPU is used to perform the graphics rendering process in the image processing method.

[0092] The display screen 194 is used to display images, videos, etc. The display screen 194 includes a display panel. In some embodiments, the electronic device may include one or N display screens 194, where N is a positive integer greater than 1. For example, the display screen 194 is used to display the preview screen and the captured images in the camera application, etc.

[0093] The electronic device can implement the shooting function through the ISP, the camera 193, the video codec, the GPU, the display screen 194, and the application processor, etc.

[0094] The camera 193 is used to capture still images or videos. In some embodiments, the electronic device may include one or N cameras 193, where N is a positive integer greater than 1. For example, after responding to the user's shooting operation, the camera 193 can be used to obtain the original image sequence.

[0095] The external memory interface 120 can be used to connect an external memory card, such as a Micro SD card, to implement the storage capacity expansion of the electronic device. The internal memory 121 can be used to store computer-executable program code, and the executable program code includes instructions. The internal memory 121 can include a storage program area and a storage data area. For example, the internal memory 121 can be used to store the executable program code in the image processing method.

[0096] The electronic device can implement the audio function through the audio module 170, the speaker 170A, the receiver 170B, the microphone 170C, the headphone interface 170D, and the application processor, etc. For example, music playback, recording, etc.

[0097] The touch sensor can be disposed on the display screen 194, and the touch sensor and the display screen 194 form a touch screen, or "touch control screen". For example, the touch sensor is used to receive the triggering operation of the user on the shooting button.

[0098] The keys 190 include a power-on key, volume keys, etc. The keys 190 can be mechanical keys or touch keys. The electronic device can receive key inputs and generate key signal inputs related to the user settings and function controls of the electronic device. In some scenarios, the electronic device can also respond to the user's operation on one or more of the keys 190 to implement the shooting operation, and the specific manner of the shooting operation in the embodiments of the present application is not limited.

[0099] The software system of the electronic device can adopt a layered architecture, an event-driven architecture, a microkernel architecture, a microservices architecture, or a cloud architecture, etc., which will not be elaborated here.

[0100] The following uses specific embodiments to elaborate in detail on the technical solution of this application and how the technical solution of this application solves the above technical problems. These several specific embodiments below can be implemented independently or in combination with each other. For the same or similar concepts or processes, they may not be elaborated in some embodiments.

[0101] Figure 3 It is a schematic diagram of the software structure of the electronic device provided by the embodiment of this application.

[0102] The layered architecture divides the system into several layers, and each layer has clear roles and divisions of labor. The layers communicate with each other through software interfaces. In some embodiments, the system is divided into five layers, from top to bottom are the application (application, APP), application framework layer (framework, FWK), hardware abstraction layer (hardware abstraction layer, HAL), driver layer, and hardware layer, etc.

[0103] The application layer may include a series of application program packages. In the embodiment of this application, the application program packages may include: camera, gallery, etc. The camera can implement the shooting of images and the display of the captured images. The gallery can also be called an album, etc., and the gallery can implement the storage and access of the captured images.

[0104] The application framework layer provides application programming interfaces (application programming interface, API) and programming frameworks for the application programs in the application layer. The application framework layer includes some predefined functions. In the embodiment of this application, the application framework layer may include a camera access interface, where the camera access interface may include camera management and camera devices. The camera access interface is used to provide application programming interfaces and programming frameworks for the camera application.

[0105] The hardware abstraction layer is an interface layer located between the application framework layer and the driver layer, providing a virtual hardware platform for the operating system. In the embodiment of this application, the hardware abstraction layer may include a camera hardware abstraction layer and a camera algorithm library.

[0106] Among them, the camera hardware abstraction layer can provide virtual hardware for camera device 1, camera device 2, or more camera devices. The camera algorithm library may include the running code and data for implementing the image processing method provided by the embodiment of this application.

[0107] The driver layer is the layer between hardware and software. The driver layer includes drivers for various hardware. The driver layer may include a camera device driver, a digital signal processor driver, and an image processor driver, etc.

[0108] Among them, the camera device driver is used to drive the sensor of the camera to collect images and drive the image signal processor to preprocess the images. The digital signal processor driver is used to drive the digital signal processor to process the images. The image processor driver is used to drive the graphics processor to process the images.

[0109] The following specifically describes the image processing method in the embodiments of the present application in combination with the above system structure:

[0110] In response to the user's operation of opening the camera application, such as the operation of clicking on the camera application icon, the camera application calls the camera access interface of the application framework layer to start the camera application, and then sends an instruction to start the camera by calling the camera device (Camera Device 1 and / or other camera devices) in the camera hardware abstraction layer. The camera hardware abstraction layer sends this instruction to the camera device driver in the kernel layer. This camera device driver can start the corresponding camera sensor and collect image optical signals through the sensor. One camera device in the camera hardware abstraction layer corresponds to one camera sensor in the hardware layer.

[0111] Then, the camera sensor can transmit the collected image optical signals to the image signal processor for preprocessing to obtain image electrical signals (raw images), and transmit the above raw images to the camera hardware abstraction layer through the camera device driver.

[0112] The camera hardware abstraction layer can send the raw images to the camera algorithm library. The camera algorithm library stores program codes for implementing the image processing method provided in the embodiments of the present application. Based on the digital signal processor and the image processor, the camera algorithm library executes the above codes to implement the process of generating the target image in the image processing method described in the embodiments of the present application.

[0113] The camera algorithm library can send the recognized raw images collected by the camera to the camera hardware abstraction layer. Then, the camera hardware abstraction layer can display them.

[0114] It can be understood that the software architecture provided in the embodiments of the present application is only an example and does not constitute a limitation to the embodiments of the present application.

[0115] Exemplarily, Figure 4 is a schematic flowchart of an image processing method provided in the embodiments of the present application. In Figure 4 the corresponding embodiment, the electronic device may include: a camera, a gallery, a camera access interface, a camera algorithm library, a camera hardware abstraction layer, and a camera device driver. The functions of any module can be referred to Figure 3 the corresponding embodiment and will not be elaborated here.

[0116] In Figure 4In the corresponding embodiment, taking the edge feature as the hair in a portrait as an example for illustration, in possible implementation manners, the edge feature may further include: the wool in a sweater, or the hair of an animal, or the branches and leaves of a plant, etc., which are not limited in the embodiments of the present application.

[0117] As Figure 4 shown, the image processing method may include the following steps:

[0118] S401. In response to a photographing operation, the camera device drives to obtain an original image.

[0119] The original image may include one or more of the following: a long frame, a short frame, or a normal frame. The exposure time of any frame in the long frame is greater than the exposure time of any frame in the normal frame, and the exposure time of any frame in the normal frame is greater than the exposure time of any frame in the short frame.

[0120] It can be understood that the original image may include at least two image frames collected at the same moment, and the exposure times of the at least two image frames are different, so that the electronic device can generate a photographed image with an HDR effect based on the image fusion of at least two image frames with different exposure times in an HDR photographing scenario.

[0121] Responding to the photographing operation includes: responding to the triggering operation of the user on the photographing button in the portrait photographing function, or responding to the triggering operation of the user on the photographing button in the HDR photographing function, or responding to the triggering operation of the user on the photographing button in the large aperture photographing function, or responding to the triggering operation of the user on the recording button in the video recording function, etc. The applicable scenarios of the image processing method are not limited in the embodiments of the present application.

[0122] S402. The camera algorithm library obtains the original image from the camera device driver.

[0123] Specifically, the camera hardware abstraction layer may obtain the original image from the camera device driver, and the camera algorithm library may obtain the original image from the camera hardware abstraction layer.

[0124] S403. The camera algorithm library processes the original image into a clear image.

[0125] The process of the camera algorithm library processing the original image into a clear image may be: the camera algorithm library may obtain at least two images with different exposure times at the same moment in the original image, and perform processes such as pre-processing on any frame image in the at least two images. Furthermore, the camera algorithm library may perform image fusion on at least two frames of images obtained at the same moment in the at least two images after pre-processing the images to obtain a clear image. The clear image may also be referred to as the first image.

[0126] It can be understood that since the clear image is obtained by fusing images with at least two different exposure times, the clear image has an HDR effect.

[0127] The pre-image processing may include one or more of the following: dead pixel correction processing, RAW domain noise reduction processing, black level correction processing, optical shadow correction processing / auto white balance processing, color interpolation processing, color correction processing, or Gamma correction processing, etc., which are not limited in the embodiments of the present application.

[0128] S404. The camera algorithm library calculates the depth of field for the clear image to obtain a blurred image corresponding to the clear image.

[0129] The blurred image may include a clear foreground and a blurred background. When the clear image is shown as Figure 1 shown in a of [figure reference], after blurring the clear image, a blurred image with a clear portrait foreground and a blurred background can be obtained. The blurred image may be the image shown in the interface of Figure 1 shown in c of [figure reference].

[0130] The camera algorithm library may determine the depth of field range based on the focus point and the depth image to obtain a depth of field image. For example, the depth at the focus point may be the first depth. In the depth of field image, the depth of field at the focus point may be 0, and the depth of field at other positions in the depth of field image may be the difference between the depth at that other position and the first depth, to obtain the foreground and the background. The foreground may include the clear range before the focus point, and the background may include the clear range after the focus point.

[0131] The focus point may be determined by the camera algorithm library based on the portrait or object captured in the clear image. For example, in the portrait shooting function of the camera, the electronic device may implement the detection of the shooting object. For example, when it is detected that the shooting object includes a portrait, the position where the portrait is located is set as the focus point to complete the focusing on the portrait, so that the portrait is clearly visible in the picture.

[0132] The depth image may include the depth information of any pixel point. The depth information may represent the distance from each point in the scene to the camera plane and can reflect the geometric shape of the visible surface in the scene. Among them, the depth information of any pixel point in the depth image may be determined based on monocular depth estimation, binocular depth estimation, or depth estimation based on deep learning, which is not limited in the embodiments of the present application.

[0133] The method of blurring processing may include: Gaussian blurring processing, or blurring processing based on a neural network, etc., which is not limited in the embodiments of the present application. For example, taking the construction method of a blurred image based on Gaussian blurring processing as an example: the camera algorithm library can perform Gaussian blurring processing with randomly varying blurring intensity on the entire clear image to obtain a Gaussian blurred image, and then extract the clear portrait foreground from the clear image according to the position of the portrait foreground in the mask image corresponding to the clear image, and replace it in the above Gaussian blurred image to obtain a blurred image.

[0134] S405. The camera algorithm library obtains a first mask image corresponding to the clear image and a first ternary image corresponding to the clear image.

[0135] There is a corresponding relationship between the pixel points in the first mask image and the pixel points in the clear image, and there is also a corresponding relationship between the pixel points in the first ternary image and the pixel points in the clear image.

[0136] Exemplarily, a target neural network (or referred to as the first model) can be pre-set in the camera algorithm library, and this target neural network is used to output a mask image and a ternary image corresponding to the image. For example, the camera algorithm library can input the clear image into the target neural network, and the target neural network can output a first mask image corresponding to the clear image and a first ternary image corresponding to the clear image. Among them, for the training process of the target neural network, reference can be made to Figure 5 the corresponding embodiments.

[0137] The target neural network may include one or more of the following branches: a shared editor, a semantic segmentation branch, a detail prediction branch, or a fusion branch, etc. Specifically, the camera algorithm library can input the clear image into the target neural network, and the shared editor can obtain the image features corresponding to the clear image and input the image features corresponding to the clear image into the semantic segmentation branch to obtain the portrait foreground. The camera algorithm library obtains the Laplacian high-frequency image corresponding to the clear image and inputs the Laplacian high-frequency image corresponding to the clear image into the detail prediction molecule to obtain the hair details in the transition area. The camera algorithm library inputs the portrait foreground and the hair details in the transition area into the fusion branch to obtain a first mask image corresponding to the clear image and a first ternary image corresponding to the clear image. Among them, the first mask image may include a portrait foreground with clear edge features, and the first ternary image may include a transition area, and the transition area may include a hair area.

[0138] Figure 5 It is a schematic diagram of the architecture of a target neural network provided by the embodiments of the present application. As Figure 5 shown, the target neural network can be a neural network based on supervised learning.

[0139] The shared encoder is used to extract image features, such as obtaining the image features of training images, and the training images can be clear images. For example, the shared encoder can adopt a pre-trained HRNet48 model on ImageNet, etc.

[0140] The semantic segmentation branch is used to obtain the approximate foreground of the human figure in the training image. For example, the semantic segmentation branch can be composed of five blocks, etc., and each block is composed of three ConvBNReLU and a bilinear upsampling module, etc. Among them, the three ConvBNReLU can include: convolution Conv, batch normalization (BN), and ReLU.

[0141] The detail prediction branch is used to predict more high-frequency information of the edges in the training image, such as obtaining high-precision hair details in the training image, etc. For example, the detail prediction branch can be composed of 3 residual blocks and convolutional layers, etc.

[0142] Specifically, the training device can use the Laplacian edge detection method to extract details from the edge region in the training image to obtain the Laplacian high-frequency map. The image features output by the shared encoder and the Laplacian high-frequency map are sent to the detail prediction branch, and the detail prediction branch can then predict a lot of high-frequency information of the edges, such as predicting the hair details in the training image.

[0143] The fusion branch is used to fuse the foreground of the human figure output by the semantic segmentation branch and the hair details output by the detail prediction branch, and output the predicted mask map and the predicted ternary map.

[0144] Exemplarily, the training data of the target neural network can include: training images, mask map ground truth, and ternary map ground truth. After the training device inputs the training images into the first neural network, the first neural network can adjust the parameters in the first neural network based on the difference between the predicted mask map and the mask map ground truth, and the difference between the predicted ternary map and the ternary map ground truth. Among them, the first neural network can also be called the source model.

[0145] For example, when the difference between the predicted mask map and the mask map ground truth satisfies the loss function (loss) loss1, and the difference between the predicted ternary map and the ternary map ground truth satisfies loss2, a trained fusion branch or a trained target nerve is obtained. Among them, loss1 can be the absolute value loss between the mask map ground truth and the predicted mask map; loss2 can be the cross-entropy loss between the ternary map ground truth and the predicted ternary map.

[0146] The predicted mask map and the preset ternary map can both be determined jointly based on the three branches of the voice segmentation branch, the detail prediction branch, and the fusion branch.

[0147] In a possible implementation, since the semantic segmentation branch performs relatively rough semantic segmentation on the image and can only obtain a general portrait foreground, while the edge of the ground truth mask is relatively accurate. Therefore, in order to enable the semantic segmentation branch to output an image consistent with the foreground contour in the ground truth mask, the training device can calculate loss4 for the ground truth mask and the output result of the semantic segmentation branch. For example, the training device can perform Gaussian blur processing on the ground truth mask to obtain a Gaussian blurred image of the ground truth mask edge (or simply referred to as a blurred mask edge image). By the difference between the blurred mask edge image and the output result of the semantic segmentation branch, the parameters in the semantic segmentation branch are adjusted, and when the blurred mask edge image and the output result of the semantic segmentation branch satisfy loss4, the training of the semantic segmentation branch is completed. Among them, loss4 can be the structural similarity loss between the blurred mask edge image and the output result of the semantic segmentation branch.

[0148] In a possible implementation, in order to enable the detail prediction branch to predict more details in the transition region containing edge features, the training device can calculate loss3 for the ternary ground truth and the output result of the detail prediction branch. For example, the training device can obtain a mask image in the ternary transition region (or simply referred to as a transition region mask image, or the third image) based on the ground truth mask and the ternary ground truth. By the difference between the transition region mask image and the output result of the detail prediction branch, the parameters in the detail prediction branch are adjusted, and when the transition region mask image and the output result of the detail prediction branch satisfy loss3, the training of the detail prediction branch is completed. Among them, loss3 can be the L1 loss that restricts the transition region mask image and the output result of the detail prediction branch. The L1 loss is also called the mean absolute error.

[0149] It can be understood that the electronic device can train a better semantic segmentation branch based on loss4, and / or train a better detail prediction branch based on loss4.

[0150] The fusion branch can combine the foreground contour in the blurred mask edge image output by the semantic segmentation branch and the clear edge features in the clear feature map output by the detail prediction branch to obtain a predicted mask image, and the predicted mask image can have clear edge features. And the fusion branch can combine the foreground contour in the blurred mask edge image output by the semantic segmentation branch and the transition region containing edge features in the clear feature map output by the detail prediction branch to obtain a predicted ternary image, and the predicted ternary image can be marked with a transition region, and the transition region can contain edge features.

[0151] Figure 5The training device described in [reference] can be the electronic device described in the embodiments of the present application, or it can also be other devices or servers that can establish a connection with the electronic device, etc. The embodiments of the present application do not limit this.

[0152] In combination with Figure 6 the corresponding embodiments, the mask image and the ternary image are illustrated by examples. Figure 6 This is a schematic diagram of an image provided by the embodiments of the present application. For example, the clear image can be seen in Figure 6 the image shown in a of [reference], the mask image can be seen in Figure 6 the image shown in b of [reference], and the ternary image can be seen in Figure 6 the image shown in c of [reference].

[0153] S406. The camera algorithm library optimizes the hair region in the blurred image by using the clear image, the first mask image, and the first ternary image to obtain the target image.

[0154] It can be understood that the camera algorithm library can obtain a first region composed of clear edge features from the clear image based on the position of the edge features in the first mask image, and fuse the first region into the blurred image (or understood as covering the first region on the blurred image) to obtain the target image. Among them, it is necessary to ensure that the edge features in the first mask image can be located in the transition region of the first ternary image.

[0155] Exemplarily, the camera algorithm library can also obtain the target image based on the generation network. Among them, the generation network can be a generative adversarial network (GAN). The generation network includes at least one generator for generating hair and at least one discriminator for judging the authenticity of the hair. The embodiments of the present application do not specifically limit the implementation details in the generation network.

[0156] Among them, the generation network can be trained based on: training images, predicted mask images, predicted ternary images, and the blurred images corresponding to the training images. The trained generation network can output a blurred image with clear edge features based on the training images, predicted mask images, predicted ternary images, and the blurred images corresponding to the training images. Among them, the predicted ternary image can be used to control the range of edge features during the training process of the generation network.

[0157] Exemplarily, the camera algorithm library can input the clear image, the first mask image, and the first ternary image into the generation network, and the generation network outputs an intermediate image. Further, the camera algorithm library can replace the region under the transition region indicated by the first ternary image in the intermediate image into the blurred image to obtain the target image. Since only the hair transition region is replaced in the final output image, the calculation range of the hair optimization module is smaller and the calculation speed is faster.

[0158] For example, the camera algorithm library may input a clear image, a first mask image, a first ternary image, and a blurred image into a generative adversarial network to output a target image, which may include clear edge features, such as clear hair strands.

[0159] After the camera algorithm library determines the target image, the camera algorithm library may store the target image in the image library through the steps shown in S407 - S408, and process the target image into a thumbnail and display the thumbnail through S409 - S411. The embodiment of the present application does not limit the sequence relationship between the above two processes.

[0160] S407. The image library obtains the target image from the camera algorithm library.

[0161] Specifically, the camera access interface (or the first interface) may obtain the thumbnail from the camera algorithm library, and then the camera may obtain the thumbnail from the camera access interface (or the first interface). Among them, the first interface (not shown in Figure 4 ) may be used to establish a data path between the image library and the camera algorithm library. The embodiment of the present application does not limit the type of the first interface.

[0162] S408. The image library stores the target image.

[0163] After S408, in response to the user's operation of opening the image library, the electronic device may display the target image.

[0164] S409. The camera algorithm library processes the target image into a thumbnail.

[0165] The camera algorithm library may process the target image into a thumbnail through thumbnail processing methods such as sampling or neural networks. The embodiment of the present application does not limit this.

[0166] S410. The camera obtains the thumbnail from the camera algorithm library.

[0167] Specifically, the camera access interface may obtain the thumbnail from the camera algorithm library, and then the camera may obtain the thumbnail from the camera access interface.

[0168] S411. The camera calls the display screen to display the thumbnail.

[0169] Exemplarily, the camera may display the thumbnail at the lower left corner of the interface shown in b in Figure 1 .

[0170] Based on this, the electronic device can obtain the mask image corresponding to the clear image and the ternary image corresponding to the clear image, determine the position of the edge feature through the mask image, determine the transition region including the edge feature through the ternary image, and then the electronic device can obtain the edge region corresponding to the position of the edge feature from the clear image based on the position of the edge feature, and fuse the edge region into the blurred image. Among them, the edge region is located within the transition region, and the blurred image can be obtained by blurring the background of the clear image.

[0171] It can be understood that Figure 4 the sequential relationship between the steps described in Figure 4 is only taken as an example and does not constitute a limitation on the embodiments of the present application.

[0172] On the basis of the Figure 4 corresponding embodiment, the embodiments of the present application provide a method for generating a three-original image, so that the transition region in the ternary image can contain edge features.

[0173] Exemplarily, Figure 7 FIG. Figure 7 is a schematic diagram of the steps for generating a ternary image provided by the embodiments of the present application. In the Figure 7 corresponding embodiment, taking the edge feature as hair strands as an example for illustration, this example does not constitute a limitation on the embodiments of the present application.

[0174] It can be understood that Figure 7 the method for generating the ternary image described in Figure 7 can be used to generate the ternary image truth value in the first neural network. For example, based on the training image and the mask image truth value, the process of generating the ternary image truth value can be referred to the description in the following Figure 7 section.

[0175] As Figure 7 shown, the method for generating the ternary image may include the following steps:

[0176] S701. The electronic device obtains the first hair strand region in the training image based on the semantic segmentation method.

[0177] For example, the electronic device can perform calculations on the training image through the Parsing-Net deep learning network to obtain the general hair strand distribution region and obtain the first hair strand region. The first hair strand region can also be referred to as the second region.

[0178] Among them, Parsing-Net is a multi-class deep learning method based on semantic segmentation, which can classify and label different regions of the human body, such as hair, face, hand, torso, leg. The labeling results obtained by such methods are usually of low accuracy. The embodiments of the present application do not specifically limit the method for obtaining the general hair strand distribution region.

[0179] S702. The electronic device performs connected component filtering on the first hair region to remove noise and obtains the second hair region.

[0180] Connected component filtering refers to finding pixel points in an image with the same or similar color values and combining them into a continuous region.

[0181] For example, the electronic device can obtain at least one connected component from the first hair region and remove connected components with an area smaller than a first value from the at least one connected component to obtain the second hair region. Among them, the second hair region can also be referred to as the third region. The first value can be a value such as 1 / 2500 of the area of the training image.

[0182] It can be understood that due to the low accuracy of the output result of the Parsing-Net method and the existence of noise, noise regions with an area smaller than 1 / 2500 of the area of the training image can be removed through connected component filtering to improve the accuracy of hair region recognition.

[0183] S703. The electronic device dilates the second hair region to obtain the third hair region.

[0184] The electronic device can perform 12 times of circular dilation with a kernel size of 6×6 on the second hair region to obtain the third hair region. The third hair region can also be referred to as the fourth region.

[0185] It can be understood that in order to improve the accuracy of hair region recognition and avoid hair regions that are not marked in the output result of the Parsing-Net method, the electronic device can perform dilation processing so that the dilated hair region can cover all hair details. Among them, in the embodiments of the present application, the number of dilations, the shape of dilation, and the size of dilation are not limited.

[0186] S704. The electronic device dilates the transition region in the ground truth mask to obtain the second mask.

[0187] The ground truth mask can be determined based on a target neural network, or the ground truth mask can also be determined by one or more of methods such as manual annotation, data synthesis, or computer graphics (CG) synthesis.

[0188] The electronic device can perform 1 time of circular dilation with a kernel size of 30×30 on the transition region in the ground truth mask to obtain the second mask region. Among them, the second mask region can also be referred to as the dilated mask. In the embodiments of the present application, the number of dilations, the shape of dilation, and the size of dilation are not limited.

[0189] The electronic device takes the intersection of the second mask region and the third hair region to obtain the transition region in the ternary graph.

[0190] It can be understood that since the transition region of the mask image can cover the entire body range, it is necessary to take the intersection of the second mask image and the third hair region, and perform a connected component filtering again to remove noise regions with an area less than 1 / 20 of the area of the entire image, etc., so as to remove the transition region distributed in the non-hair region and obtain the transition region in the ternary graph. The transition region in the ternary graph can also be referred to as the first transition region.

[0191] S706. The electronic device sets the region where alpha = 255 in the mask image truth value in the foreground of the ternary graph to obtain the ternary graph truth value. The electronic device can fuse the region where alpha = 255 in the mask image truth value and the third hair region to obtain the ternary graph truth value.

[0192] Based on this, the electronic device can obtain the ternary graph truth value with the transition region restricted to the hair position by fusing the mask region and the hair region.

[0193] In a possible implementation manner, the embodiment of the present application can also be based on Figure 7 the ternary graph generation method described in Figure 7 to obtain the first ternary graph corresponding to the clear image. The specific process is similar to that in

[0194] and will not be elaborated here. Figure 4 Based on the corresponding embodiment, in order to clearly illustrate, the embodiment of the present application also provides an image processing method. Figure 8 This is a schematic flowchart of another image processing method provided by the embodiment of the present application. In Figure 8 the corresponding embodiment, taking the edge feature as the hair is used as an example for illustration, and this example does not limit the embodiment of the present application.

[0195] The camera algorithm library of the electronic device may include: a depth calculation module, a hair matte extraction module, and a hair optimization module. The depth calculation module is used to obtain a blurred image with a clear foreground and a blurred background. The hair matte extraction module is used to obtain a first mask image containing hair and a first ternary graph containing hair. The hair optimization module is used to obtain a target image with clear hair by using the clear image, the blurred image, the mask image, and the ternary graph.

[0196] Among them, the implementation details of the depth calculation module can be referred to the description in S404, the implementation details of the hair matte extraction module can be referred to the description in S405, and the implementation details of the hair optimization module can be referred to the description in S406, and will not be elaborated here.

[0197] As described above in combination with Figures 4 - 8, the method provided in the embodiments of the present application has been described. Next, the apparatus for executing the above method provided in the embodiments of the present application will be described. As Figure 9 shown, Figure 9 FIG. is a schematic structural diagram of an image processing apparatus provided in an embodiment of the present application. The image processing apparatus may be an electronic device in the embodiments of the present application, or a chip or a chip system in the electronic device.

[0198] As Figure 9 shown, the image processing apparatus 900 may be used in a communication device, a circuit, a hardware component, or a chip. The image processing apparatus 900 includes: a display unit 901 and a processing unit 902. Among them, the display unit 901 is used to support the display steps performed by the display method; the processing unit 902 is used to support the image processing apparatus 900 to perform information processing steps.

[0199] In a possible implementation manner, the image processing apparatus 900 may further include a communication unit 903. The communication unit 903 is used to support the image processing apparatus 900 to perform steps such as receiving or sending messages.

[0200] The image processing apparatuses described in the embodiments of the present application may all include Figure 9 the units described in the corresponding embodiments.

[0201] Specifically, the processing unit 902 and the display unit 901 may be integrated together, and communication may occur between the processing unit 902 and the display unit 901.

[0202] In a possible implementation manner, the image processing apparatus 900 may further include: a storage unit 904. Among them, the storage unit 904 may include one or more memories. The memory may be a device or a circuit in one or more devices for storing programs or data.

[0203] The storage unit 904 may exist independently and be connected to the processing unit 902 through a communication bus. The storage unit 904 may also be integrated with the processing unit 902.

[0204] Taking the image processing device 900 as an example of the chip or chip system of the electronic device in the embodiments of the present application, the storage unit 904 may store computer-executable instructions of the method of the electronic device, so that the processing unit 902 executes the method of the electronic device in the above embodiments. The storage unit 904 may be a register, a cache, or a random access memory (RAM), etc., and the storage unit 904 may be integrated with the processing unit 902. The storage unit 904 may be a read-only memory (ROM) or other types of static storage devices that can store static information and instructions, and the storage unit 904 may be independent of the processing unit 902.

[0205] In a possible implementation, the image processing device 900 may further include: a communication unit 903. Among them, the communication unit 903 is used to support the image processing device 900 to interact with other devices. Exemplarily, when the image processing device 900 is an electronic device, the communication unit 903 may be a communication interface or an interface circuit. When the image processing device 900 is a chip or chip system inside an electronic device, the communication unit 903 may be a communication interface. For example, the communication interface may be an input / output interface, a pin, or a circuit, etc.

[0206] The device in this embodiment can correspondingly be used to execute the steps executed in the above method embodiment, and its implementation principle and technical effect are similar, and will not be elaborated here.

[0207] Figure 10 This is a schematic diagram of the hardware structure of another electronic device provided by the embodiments of the present application.

[0208] The electronic device includes a processor 1001, a communication line 1004, and at least one communication interface ( Figure 10 exemplarily described by taking the communication interface 1003 as an example).

[0209] The processor 1001 may be a general-purpose central processing unit (CPU), a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits for controlling the execution of the program of the solution of the present application.

[0210] The communication line 1004 may include a circuit for transmitting information between the above components.

[0211] The communication interface 1003 uses any device such as a transceiver for communicating with other devices or communication networks, such as Ethernet, wireless local area networks (WLAN), etc.

[0212] Optionally, the electronic device may further include a memory 1002.

[0213] The memory 1002 can be a read-only memory (ROM) or other types of static storage devices that can store static information and instructions, a random access memory (RAM) or other types of dynamic storage devices that can store information and instructions, or it can also be an electrically erasable programmable read-only memory (EEPROM), a compact disc read-only memory (CD-ROM) or other optical disc storage, optical disc storage (including compact discs, laser discs, optical discs, digital versatile discs, Blu-ray discs, etc.), magnetic disk storage media or other magnetic storage devices, or any other medium that can be used to carry or store the desired program code in the form of instructions or data structures and can be accessed by a computer, but is not limited thereto. The memory can exist independently and be connected to the processor through a communication line 1004. The memory can also be integrated with the processor.

[0214] Among them, the memory 1002 is used to store computer execution instructions for implementing the solution of this application, and is controlled by the processor 1001 to execute. The processor 1001 is used to execute the computer execution instructions stored in the memory 1002, so as to implement the method provided by the embodiments of this application.

[0215] Optionally, the computer execution instructions in the embodiments of this application may also be referred to as application code, and the embodiments of this application do not make specific limitations thereon.

[0216] In a specific implementation, as an embodiment, the processor 1001 may include one or more CPUs, such as Figure 10 CPU0 and CPU1 in

[0217] In a specific implementation, as an embodiment, the electronic device may include multiple processors, such as Figure 10The processors 1001 and 1005 therein. Each of these processors can be a single-CPU processor or a multi-CPU processor. The processors herein can refer to one or more devices, circuits, and / or processing cores for processing data (such as computer program instructions).

[0218] In the above embodiments, the instructions stored in the memory for the processor to execute can be implemented in the form of a computer program product. Among them, the computer program product can be pre-written in the memory or downloaded and installed in the memory in the form of software.

[0219] The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions according to the embodiments of the present application are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions can be transmitted from a website, computer, server, or data center to another website, computer, server, or data center in a wired manner (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or wirelessly (such as infrared, wireless, microwave, etc.). The computer-readable storage medium can be any available medium that can be stored by a computer or a data storage device such as a server, data center, etc. that includes one or more available media integrated. For example, the available medium can include magnetic media (such as floppy disks, hard disks, or magnetic tapes), optical media (such as digital versatile discs (DVDs)), or semiconductor media (such as solid state disks (SSDs)), etc.

[0220] The embodiments of the present application also provide a computer-readable storage medium. The methods described in the above embodiments can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. The computer-readable medium can include a computer storage medium and a communication medium, and can also include any medium that can transfer a computer program from one place to another. The storage medium can be any target medium accessible by a computer.

[0221] As a possible design, a computer-readable medium may include a compact disc read-only memory (CD-ROM), RAM, ROM, EEPROM, or other optical disc storage; the computer-readable medium may include magnetic disk storage or other magnetic disk storage devices. Moreover, any connecting wire may also be appropriately referred to as a computer-readable medium. For example, if software is transmitted from a website, server, or other remote source using coaxial cable, fiber optic cable, twisted pair, DSL, or wireless technologies such as infrared, radio, and microwave, then the coaxial cable, fiber optic cable, twisted pair, DSL, or wireless technologies such as infrared, radio, and microwave are included in the definition of the medium. As used herein, magnetic disks and optical discs include optical discs (CDs), laser discs, optical discs, digital versatile discs (DVDs), floppy disks, and Blu-ray discs, where magnetic disks typically reproduce data magnetically, while optical discs use lasers to optically reproduce data.

[0222] The above combinations should also be included within the scope of computer-readable media. The above are only specific embodiments of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention can easily think of changes or substitutions, which should all be covered within the protection scope of the present invention. Therefore, the protection scope of the present invention should be subject to the protection scope of the claims.

[0223] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in this application are all information and data that have been authorized by the user or fully authorized by all parties. Moreover, the collection, use, and processing of relevant data need to comply with relevant laws, regulations, and standards, and corresponding operation entrances are provided for users to choose to authorize or refuse.

Claims

1. An image processing method, characterized in that, the method includes: collecting a first image in response to a photographing operation; obtaining a blurred image corresponding to the first image, the positions of edge features in the first image, and a transition region formed by the edge features in the first image; wherein, the blurred image is obtained by blurring the background in the first image; using the positions of the edge features to determine a first region formed by the edge features from the first image, and fusing the first region into the blurred image to obtain a target image; wherein, the first region is within the transition region.

2. The method according to claim 1, characterized in that, before obtaining the positions of edge features in the first image and the transition region formed by the edge features in the first image, the method further includes: obtaining a mask image corresponding to the first image and a ternary image corresponding to the first image, wherein the mask image corresponding to the first image includes the positions of the edge features, and the ternary image corresponding to the first image includes the transition region.

3. The method according to claim 1 or 2, characterized in that, obtaining the mask image corresponding to the first image and the ternary image corresponding to the first image includes: inputting the first image into a first model to output the mask image corresponding to the first image and the ternary image corresponding to the first image; wherein, the first model includes: a shared editor, a semantic segmentation branch, a detail prediction branch, and a fusion branch, the shared editor is used to extract image features of the first image, the semantic segmentation branch is used to perform semantic segmentation on the first image and obtain the foreground in the first image, the detail prediction branch is used to obtain the edge features in the first image, and the fusion branch is used to fuse the foreground in the first image and the edge features in the first image to obtain the mask image corresponding to the first image and the ternary image corresponding to the first image.

4. The method according to claim 3, characterized in that, the method further includes: the first model is trained in the following manner: inputting a training image into the shared editor to obtain image features of the training image; inputting the image features of the training image into the semantic segmentation branch to obtain the output result of the semantic segmentation branch; obtaining a Laplacian high-frequency map corresponding to the training image, inputting the Laplacian high-frequency map into the detail prediction branch to obtain the output result of the detail prediction branch; inputting the output result of the semantic segmentation branch and the output result of the detail prediction branch into the fusion branch to obtain a predicted mask image and a predicted ternary image; when the difference between the predicted mask image and the mask image ground truth satisfies a first loss function, and the difference between the predicted ternary image and the ternary image ground truth satisfies a second loss function, the trained first model is obtained.

5. The method according to claim 4, characterized in that, the method further includes: Perform Gaussian blur processing on the edges of the ground truth of the mask image to obtain a second image; When the difference between the second image and the output result of the semantic segmentation branch satisfies the fourth loss function, the trained semantic segmentation branch is obtained.

6. The method according to claim 4 or 5, wherein, the method further includes: According to the transition region in the ternary ground truth, obtain the image within the transition region from the ground truth of the mask image to obtain a third image; When the difference between the third image and the output result of the detail prediction branch satisfies the third loss function, the trained detail prediction branch is obtained.

7. The method according to any one of claims 4-6, wherein, the method further includes: Obtain a second region containing edge features from the training image based on a semantic segmentation method; Perform connected component filtering on the second region to obtain a third region; Perform dilation on the third region to obtain a fourth region; Perform dilation on the ground truth of the mask image to obtain a dilated mask image; Obtain the intersection of the dilated mask image and the fourth region to obtain a first transition region; Determine the ternary ground truth based on the first transition region and the foreground in the dilated mask image.

8. The method according to any one of claims 1-7, wherein, the edge features include one or more of the following: the hair of a person, the wool in a sweater, or the hair of an animal, or the branches and leaves of a plant.

9. The method according to any one of claims 1-8, wherein, the collecting a first image in response to a photographing operation includes: In response to the photographing operation, obtain at least two images with different exposure times, and perform image fusion on the at least two images with different exposure times to obtain the first image.

10. An electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein, when the processor executes the computer program, the electronic device executes the method according to any one of claims 1-9.

11. A computer-readable storage medium, the computer-readable storage medium stores a computer program, wherein, when the computer program is executed by a processor, the computer executes the method according to any one of claims 1-9.

12. A computer program product, wherein, includes a computer program, when the computer program is run, the computer executes the method according to any one of claims 1-9.

13. A chip, wherein, includes: a processor for reading instructions stored in a storage memory, and when the processor executes the instructions, the chip implements the method according to any one of claims 1-9 above.

Citation Information

Patent Citations

  • Image background blurring method and device, storage medium and electronic device

    CN110009556A

  • Portrait extraction method, portrait extraction device and mobile terminal

    CN111507994A

  • Image processing method and related equipment

    CN111539960A

  • Image processing method and electronic equipment

    CN112116624A

  • Face beautifying method and device, electronic equipment and storage medium

    CN112561822A