Image processing method and device

By acquiring the images of the binocular camera, using semantic segmentation and deep learning models to determine the edge area of the portrait, filtering and image fusion, and generating a more accurate parallax map, it solves the problem of false or misleading electronic devices in the blurring process, and improves the quality of photography.

CN120339149APending Publication Date: 2025-07-18HONOR DEVICE CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202410042880.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-01-10
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

In the prior art, electronic devices are prone to false or missed when performing image blurring processing, resulting in a decrease in the quality of taking pictures.

Method used

By acquiring the first and second images taken by the binocular camera, the portrait edge area is determined using semantic segmentation and deep learning models, filtering and image fusion are performed, more accurate parallax maps are generated, and then blurring is performed.

Benefits of technology

It improves the accuracy of image blurring processing, reduces the edge error or background leakage caused by parallax estimation errors, and improves the quality of photography.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120339149A_ABST
    Figure CN120339149A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides an image processing method and device, and relates to the technical field of terminals. The method comprises the steps that a first disparity map and an edge area in a first image are acquired, the first disparity map is determined based on the visual difference between the first image and a second image, and the edge area is an area formed by portrait features in the first image; obtaining a first area with the same position as the edge area from the first disparity map, and filtering the first area to obtain a second area; fusing the second region into the first disparity map to obtain a second disparity map; and performing blurring processing on the first image for the second disparity map to obtain a target image. Therefore, as the edge of the portrait in the second disparity map is clear, the first image is blurred based on the second disparity map, the situation that the edge is false or the background is missed due to the fact that disparity estimation is wrong can be reduced, and the photographing quality is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of terminal technologies, and in particular, to an image processing method and apparatus. Background Art

[0002] With the popularization and development of the Internet, people's functional requirements for electronic devices have become more and more diverse. For example, an electronic device can not only support a shooting function, but also support blurring processing of the captured image, making the captured image after blurring processing more vivid.

[0003] Generally, in order to improve image quality, an electronic device can use a binocular camera to determine a disparity map, further estimate depth information, and perform blurring processing on the captured image based on the depth information to obtain a blurred image.

[0004] However, the above image processing method may cause the blurred image to have false blurring or missed blurring. Summary of the Invention

[0005] Embodiments of this application provide an image processing method and apparatus, which are applied to the field of terminal technologies and are used to improve the accuracy of image processing.

[0006] In a first aspect, embodiments of this application propose an image processing method. The method includes: in response to a photographing operation, obtaining a first image and a second image, where the first image and the second image are images obtained based on a binocular camera; obtaining a first disparity map and an edge region in the first image, where the first disparity map is determined based on the visual difference between the first image and the second image, and the edge region is a region composed of human portrait features in the first image; obtaining a first region in the first disparity map that has the same position as the edge region, performing filtering processing on the first region to obtain a second region; fusing the second region into the first disparity map to obtain a second disparity map; and performing blurring processing on the first image using the second disparity map to obtain a target image.

[0007] The edge region may be the first edge region described in the embodiments of this application, the first region may be the second edge region described in the embodiments of this application, and the second region may be the third edge region described in the embodiments of this application.

[0008] It can be understood that the electronic device can obtain the first image and the second image based on the binocular camera. By performing semantic segmentation on the first image, an edge region containing the edges of the human figure is determined. In the first disparity map, a first region with the same position as the edge region is obtained. By performing filtering processing on the first region, the clarity of the human figure edges in the first region is improved. Furthermore, after fusing the first region and the first disparity map, a second disparity map is obtained. In this way, since the human figure edges in the second disparity map are relatively clear, performing blurring processing on the first image based on the second disparity map can reduce the situation of incorrect edge blurring or background leakage blurring caused by incorrect disparity estimation, and improve the photo-taking quality.

[0009] In a possible implementation manner, before obtaining the first disparity map and the edge region in the first image, the method further includes: performing semantic segmentation on the first image to obtain a third image with a human figure region; performing dilation processing on the third image to obtain a fourth image; using the third image and the fourth image to determine the edge region.

[0010] The third image may be the human figure segmentation map described in the embodiments of the present application, and the fourth image may be the dilation map described in the embodiments of the present application.

[0011] The electronic device can make the edge region include all the human figure edges in the human figure segmentation map through dilation processing. In this way, the subsequent camera algorithm library can obtain a first region with the same position as the edge region from the first disparity map, and perform filtering processing on the first region to improve the accuracy of the human figure edges in the disparity map.

[0012] In a possible implementation manner, performing filtering processing on the first region to obtain a second region includes: performing minimum filtering and / or Gaussian filtering processing on the first region to obtain the second region.

[0013] Since the human figure segmentation map cannot ensure edge fitting, the electronic device can obtain a relatively fitting edge region through filtering processing. In the disparity map, the disparity of an object closer to the binocular camera is larger. Therefore, by performing minimum filtering to obtain the minimum value of the disparity map within a certain region 1, the minimum value of the disparity map can be the disparity map of the background in region 1, so as to obtain clearer and more fitting human figure edges. Gaussian filtering can be used to smooth the human figure edges, making the transition of the edge region relatively smooth.

[0014] In a possible implementation manner, obtaining the first disparity map includes: inputting the first image and the second image into a deep learning model, and outputting the first disparity map. The deep learning model is used to determine the first disparity map by using the human figure segmentation information of the first image, the first image, and the matching relationship between the first image and the second image.

[0015] It can be understood that since the deep learning module can utilize the human segmentation information to optimize the generation of the disparity map, the first disparity map can have a relatively accurate human portrait edge, improving the accuracy of the disparity map.

[0016] In a possible implementation, the deep learning model includes: a correlation map calculation module and a disparity estimation module. The correlation map calculation module is used to perform feature matching on the first image and the second image to obtain a matching relationship. The disparity estimation module is used to determine the first disparity map by using the human segmentation information and the matching relationship.

[0017] In a possible implementation, the deep learning model further includes: an atrous spatial pyramid pooling (ASPP) module, a feature encoding module, and a feature decoding module. The feature encoding module is used to perform feature extraction on the first image and the second image respectively to obtain a first feature map corresponding to the first image and a second feature map corresponding to the second image. The ASPP module is used to perform feature extraction on the first feature map to obtain a third feature map. The feature decoding module is used to decode the third feature map into a fourth feature map, and the fourth feature map contains human segmentation information. The correlation calculation module is used to perform feature matching on the first feature map and the second feature map to obtain a matching relationship.

[0018] As Figure 6 described in, the first feature map can be feature Figure 1 , and the second image can be feature Figure 6 , the third feature map can be feature Figure 2 , and the fourth feature map can be feature Figure 4 .

[0019] It can be understood that the ASPP module can implement feature extraction of images of any size through multiple atrous convolutions with different sampling rates, and a higher-precision semantic recognition result can be obtained based on the output result of the ASPP module.

[0020] In a possible implementation, the deep learning model further includes: a first splicing module and a second splicing module. The method further includes: the first splicing module is used to splice the third feature map and the first feature map to obtain a fifth feature map; the second splicing module is used to splice the fourth feature map and the first feature map to obtain a sixth feature map, and the sixth feature map contains human segmentation information.

[0021] As Figure 6 described in, the first splicing module can be splicing module 603, and the second splicing module can be splicing module 605. The fifth feature map can be feature Figure 3 , and the sixth feature map can be feature Figure 5 .

[0022] It is understandable that there may be cases where some information of the image is lost or added after each image processing. Therefore, the higher the complexity of the image can be increased through image stitching, the higher the accuracy of subsequent processing based on the stitched image can be improved.

[0023] In a second aspect, an embodiment of the present application provides an image processing method, including: obtaining a first image and a second image, where the first image and the second image are images obtained based on a binocular camera; inputting the first image and the second image into a deep learning model to output a first disparity map, and the deep learning model is used to determine the first disparity map by using the human segmentation information of the first image, the first image, and the matching relationship between the first image and the second image.

[0024] In a possible implementation manner, the deep learning model includes: a correlation map calculation module and a disparity estimation module. The correlation map calculation module is used to perform feature matching on the first image and the second image to obtain a matching relationship, and the disparity estimation module is used to determine the first disparity map by using the human segmentation information and the matching relationship.

[0025] In a possible implementation manner, the deep learning model further includes: an atrous spatial pyramid pooling (ASPP) module, a feature encoding module, and a feature decoding module; the feature encoding module is used to perform feature extraction on the first image and the second image respectively to obtain a first feature map corresponding to the first image and a second feature map corresponding to the second image, the ASPP module is used to perform feature extraction on the first feature map to obtain a third feature map, the feature decoding module is used to decode the third feature map into a fourth feature map, the fourth feature map contains human segmentation information, and the correlation calculation module is used to perform feature matching on the first feature map and the second feature map to obtain a matching relationship.

[0026] In a possible implementation manner, the deep learning model further includes: a first stitching module and a second stitching module. The method further includes: the first stitching module is used to stitch the third feature map and the first feature map to obtain a fifth feature map; the second stitching module is used to stitch the fourth feature map and the first feature map to obtain a sixth feature map, and the sixth feature map contains human segmentation information.

[0027] In a third aspect, an embodiment of the present application provides an image processing apparatus, which may be an electronic device, or a chip or a chip system within the electronic device. The image processing apparatus may include an acquisition unit and a processing unit. The acquisition unit is configured to perform the step of data acquisition, so that the electronic device implements an image processing method described in the first aspect or any one of the possible implementation manners of the first aspect. When the image processing apparatus is an electronic device, the processing unit may be a processor. The image processing apparatus may further include a storage unit, which may be a memory. The storage unit is configured to store instructions, and the processing unit executes the instructions stored in the storage unit, so that the electronic device implements an image processing method described in the first aspect or any one of the possible implementation manners of the first aspect. When the image processing apparatus is a chip or a chip system within the electronic device, the processing unit may be a processor. The processing unit executes the instructions stored in the storage unit, so that the electronic device implements an image processing method described in the first aspect or any one of the possible implementation manners of the first aspect. The storage unit may be a storage unit within the chip (e.g., registers, caches, etc.), or a storage unit outside the chip within the electronic device (e.g., read-only memory, random access memory, etc.).

[0028] Specifically, the image processing apparatus relates to an acquisition unit and a processing unit. In response to a photographing operation, the acquisition unit is configured to acquire a first image and a second image, where the first image and the second image are images acquired based on a binocular camera; the acquisition unit is further configured to acquire a first disparity map and an edge region in the first image, where the first disparity map is determined based on the visual difference between the first image and the second image, and the edge region is a region formed by portrait features in the first image; the processing unit is configured to acquire a first region in the first disparity map that has the same position as the edge region, perform filtering processing on the first region to obtain a second region; the processing unit is further configured to fuse the second region into the first disparity map to obtain a second disparity map; and perform blurring processing on the first image using the second disparity map to obtain a target image.

[0029] In a fourth aspect, an embodiment of the present application provides an image processing apparatus, which may be an electronic device, or a chip or a chip system within the electronic device.

[0030] Specifically, the image processing apparatus relates to an acquisition unit and a processing unit. The acquisition unit is configured to acquire a first image and a second image, where the first image and the second image are images acquired based on a binocular camera; the processing unit is configured to input the first image and the second image into a deep learning model, and output a first disparity map, where the deep learning model is used to determine the first disparity map using the portrait segmentation information of the first image and the matching relationship between the first image and the second image.

[0031] In a fifth aspect, an embodiment of the present application provides an electronic device, which includes: one or more processors and a memory; the memory is coupled to the one or more processors, and the memory is used to store computer program code. The computer program code includes computer instructions, and the one or more processors call the computer instructions to enable the electronic device to execute the method described in the first aspect or any possible implementation manner of the first aspect, or execute the method described in the second aspect or any possible implementation manner of the second aspect.

[0032] In a sixth aspect, an embodiment of the present application provides a computer-readable storage medium, which includes computer instructions. When the computer instructions run on an electronic device, the electronic device is enabled to execute the method described in the first aspect or any possible implementation manner of the first aspect, or execute the method described in the second aspect or any possible implementation manner of the second aspect.

[0033] In a seventh aspect, an embodiment of the present application provides a computer program product including a computer program. When the computer program product includes computer program code, and when the computer program code runs on an electronic device, the electronic device is enabled to execute the method described in the first aspect or any possible implementation manner of the first aspect, or execute the method described in the second aspect or any possible implementation manner of the second aspect.

[0034] In an eighth aspect, the present application provides a chip system, which is applied to an electronic device. The chip system includes one or more processors, and the one or more processors are used to call computer instructions to enable the electronic device to execute the method described in the first aspect or any possible implementation manner of the first aspect, or execute the method described in the second aspect or any possible implementation manner of the second aspect.

[0035] In a possible implementation, the chip system described above in the present application further includes at least one memory, and instructions are stored in the at least one memory. The memory may be an internal storage unit of the chip system, for example, a register, a cache, etc., or may be a storage unit of the chip system (for example, a read-only memory, a random access memory, etc.).

[0036] It should be understood that the technical solutions of the second to eighth aspects of the present application correspond to those of the first aspect of the present application, and the beneficial effects obtained by each aspect and the corresponding feasible implementation manners are similar, and will not be elaborated herein. BRIEF DESCRIPTION OF THE DRAWINGS

[0037] Figure 1 It is a schematic diagram of a scenario provided by an embodiment of the present application;

[0038] Figure 2Schematic diagram of the hardware structure of an electronic device provided by an embodiment of the present application;

[0039] Figure 3 Schematic diagram of the software structure of an electronic device provided by an embodiment of the present application;

[0040] Figure 4 Schematic diagram of the process of an image processing method provided by an embodiment of the present application;

[0041] Figure 5 Schematic diagram of an image provided by an embodiment of the present application;

[0042] Figure 6 Schematic diagram of a deep learning model provided by an embodiment of the present application;

[0043] Figure 7 Schematic diagram of the structure of a feature encoding module provided by an embodiment of the present application;

[0044] Figure 8 Schematic diagram of the structure of an ASPP module provided by an embodiment of the present application;

[0045] Figure 9 Schematic diagram of the structure of a feature decoding module provided by an embodiment of the present application;

[0046] Figure 10 Schematic diagram of the structure of a portrait segmentation module provided by an embodiment of the present application;

[0047] Figure 11 Schematic diagram of the steps of an image processing method provided by an embodiment of the present application;

[0048] Figure 12 Schematic diagram of the structure of an image processing device provided by an embodiment of the present application;

[0049] Figure 13 Schematic diagram of the hardware structure of another electronic device provided by an embodiment of the present application. Detailed implementation manners

[0050] For the convenience of clearly describing the technical solutions of the embodiments of the present application, the following briefly introduces some terms and technologies involved in the embodiments of the present application:

[0051] 1. Alpha mask image (or simply referred to as mask image)

[0052] The mask image can be an image generated by performing occlusion processing on an image (all or part of it), and the mask image can be used to extract the region of interest. Taking the data format of the mask image as 8-bit as an example, the alpha value of each pixel point in the mask image is illustrated. In the mask image, the alpha value of the pixel points is distributed between 0 and 255. The foreground region of interest can be set to alpha = 255, and the background region can be set to alpha = 0.

[0053] In the embodiments of the present application, the electronic device can determine the portrait region in the image through semantic segmentation of the image, and set the alpha of the portrait region to 255 to obtain the portrait segmentation image.

[0054] 2. Disparity Map

[0055] The disparity map can be used to reflect the visual difference between two images obtained based on a binocular camera. Binocular stereo vision fuses the images obtained by both eyes and observes the differences between them to obtain an obvious sense of depth. The electronic device can establish the correspondence between feature points and correspond the image points of the same physical point in space in different images. The visual difference between feature points can form a disparity map.

[0056] It can be understood that the closer the object is to the two eyes, the greater the disparity. For example, place the finger at different distances from the eyes and alternately open and close the left eye and open and close the right eye. It can be found that at different distances of the finger, the disparity is also different, and the closer the distance, the greater the disparity.

[0057] There is a correspondence between disparity and depth. After the electronic device calculates the disparity map, it can convert the disparity map into a depth map.

[0058] 3. Atrous Spatial Pyramid Pooling (ASPP) Module

[0059] The ASPP module is a module used to implement image segmentation tasks, aiming to solve the problem of insufficient spatial context information in semantic segmentation. The ASPP module is based on the idea of atrous convolution. Atrous convolution is a technique that can increase the receptive field without increasing the network parameters, which can help the model obtain image information in a larger range. The ASPP module uses multiple atrous convolutions. Atrous convolutions with different sampling rates can capture local information of different scales, and finally obtain feature maps with different receptive fields.

[0060] Among them, the receptive field can be: the size of the area on the input image to which the pixel points on the feature map output by each layer of the convolutional neural network are mapped back. Or it can be understood as the size of a point on the feature map relative to the original image, which is also the area of the input image that the features of the convolutional neural network can see.

[0061] 4. Other terms

[0062] In the embodiments of the present application, terms such as "first" and "second" are used to distinguish identical or similar items with basically the same functions and roles. For example, the first chip and the second chip are only used to distinguish different chips and do not limit their order. Those skilled in the art can understand that terms such as "first" and "second" do not limit the quantity and execution order, and "first", "second", etc. do not necessarily mean different.

[0063] It should be noted that in the embodiments of the present application, words such as "exemplary" or "for example" are used to indicate examples, illustrations or explanations. Any embodiment or design solution described as "exemplary" or "for example" in the present application should not be construed as being more preferred or having more advantages than other embodiments or design solutions. Rather, the use of words such as "exemplary" or "for example" is intended to present relevant concepts in a specific manner.

[0064] In the embodiments of the present application, "at least one" means one or more, and "a plurality" means two or more. "And / or" describes the association relationship of associated objects, indicating that there can be three relationships. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone, where A and B can be singular or plural. The character " / " generally represents an "or" relationship between the associated objects before and after. "At least one (item)" or its similar expression refers to any combination of these items, including any combination of single item (item) or plural items (items). For example, at least one (item) of a, b, or c can represent: a, b, c, a - b, a - c, b - c, or a - b - c, where a, b, c can be single or multiple.

[0065] 5. Electronic device

[0066] The electronic device according to the embodiments of the present application may include a handheld device with a display function, a vehicle-mounted device, etc. For example, some electronic devices may be: mobile phone, tablet computer, handheld computer, laptop computer, mobile internet device (MID), wearable device, virtual reality (VR) device, augmented reality (AR) device, etc. The embodiments of the present application are not limited thereto.

[0067] As an example rather than a limitation, in the embodiments of the present application, the electronic device may also be a wearable device, such as glasses, gloves, watches, clothing, shoes, etc.

[0068] The electronic device in the embodiments of the present application may also be referred to as: terminal device, user equipment (UE), mobile station (MS), mobile terminal (MT), access terminal, user unit, user station, mobile station, mobile platform, remote station, remote terminal, mobile device, user terminal, terminal, wireless communication device, user agent or user device, etc.

[0069] In daily photography, keeping the foreground clear and blurring the background is a common photography method. The electronic device can improve the photography quality by blurring the photographed image. During the blurring process, the electronic device can rely on binocular stereo parallax or depth estimation, etc., to determine the portrait area and the background area, and then blur the background area to obtain the blurring result. However, when there is an error in the parallax estimation in the main portrait area, it will cause the blurring result to have the portrait edge blurred by mistake or the background area not blurred.

[0070] The following Figure 1 is a schematic illustration of the photographing scene. Exemplarily, Figure 1 is a schematic diagram of a scene provided by the embodiments of the present application. In Figure 1 the corresponding embodiment, taking the electronic device as a mobile phone as an example for illustration, this example does not limit the embodiments of the present application.

[0071] As Figure 1 shown in a of, the scene may include a person in the foreground, and the sun and sun rays in the background, etc. At least one hair area may be displayed around the person's head, such as hair area 101, and hair area 102, etc.

[0072] In response to the user turning on the portrait photography function in the camera, the electronic device may display as Figure 1The interface shown as b in [the figure] can be the preview interface in the portrait photography function. The interface may include: a capture button 103 and a preview screen 104. The preview image displayed in the preview screen 104 can be obtained by the electronic device performing a blurring process on the captured image. In possible implementation manners, the blurred preview image may not be displayed in the preview screen either, and the embodiments of the present application do not make any limitations in this regard.

[0073] In response to a triggering operation by the user on the capture button 103, the electronic device Figure 1 captures the scene shown as a in [the figure] and obtains a captured image 105 in the interface shown as c in [the figure]. Among them, the captured image 105 can be obtained by the electronic device performing a blurring process on the captured image. Figure 1 The captured image 105 may include: a clear hair region 101', a blurred hair region 102', and a clear partial background region 106, etc. Refer to

[0074] the scene in a in [the figure] and Figure 1 the captured image shown as c in [the figure]. The hair region 101 can be a part of the foreground region, so the hair region 101' can be a clear region. The hair region 102 can be a part of the foreground. Due to an error in the electronic device's parallax estimation, the hair region 102 is misjudged as the background, resulting in the hair region 102' being blurred. The sun and its rays can be the background. Due to an error in the electronic device's parallax estimation, some of the sun's rays are misjudged as a part of the portrait, resulting in some of the background region 106 not being blurred. Figure 1 It can be understood that an incorrect parallax estimation will cause the blurring processing result to have a situation where the portrait edge is blurred by mistake or the background region is not blurred, affecting the photo quality.

[0075] In view of this, the embodiments of the present application provide an image processing method, enabling the electronic device to obtain a first image and a second image based on a binocular camera. By performing semantic segmentation on the first image, an edge region including the portrait edge is determined. A first region with the same position as the edge region is obtained in the first disparity map. By performing a filtering process on the first region, the clarity of the portrait edge in the first region is improved. Then, the first region and the first disparity map are fused to obtain a second disparity map. In this way, since the portrait edge in the second disparity map is relatively clear, performing a blurring process on the first image based on the second disparity map can reduce the situation of incorrect edge blurring or missed background blurring caused by incorrect parallax estimation, improving the photo quality.

[0076] Therefore, in order to better understand the embodiments of the present application, the structure of the electronic device in the embodiments of the present application will be introduced below. Exemplarily,

[0077] Therefore, in order to better understand the embodiments of the present application, the structure of the electronic device in the embodiments of the present application will be introduced below. Exemplarily, Figure 2A schematic diagram of the hardware structure of an electronic device provided by an embodiment of the present application.

[0078] The electronic device may include a processor 110, an external memory interface 120, an internal memory 121, a universal serial bus (USB) interface 130, a charging management module 140, a power management module 141, an antenna 1, an antenna 2, a mobile communication module 150, a wireless communication module 160, an audio module 170, a speaker 170A, a receiver 170B, a microphone 170C, a headphone interface 170D, a sensor module 180, a button 190, an indicator 192, a camera 193, and a display screen 194, etc.

[0079] Among them, the sensor module 180 may include one or more of the following: a pressure sensor, a gyroscope sensor, a barometric pressure sensor, a magnetic sensor, an acceleration sensor, a distance sensor, a proximity light sensor, a fingerprint sensor, a temperature sensor, a touch sensor, an ambient light sensor, or a bone conduction sensor ( Figure 2 not shown in the figure), etc., and the embodiments of the present application do not make specific limitations thereto.

[0080] It can be understood that the structure schematically shown in the embodiments of the present application does not constitute a specific limitation to the electronic device. In other embodiments of the present application, the electronic device may include more or fewer components than shown in the figure, or combine certain components, or split certain components, or have different component arrangements. The illustrated components may be implemented in hardware, software, or a combination of software and hardware.

[0081] The processor 110 may include one or more processing units. Among them, different processing units may be independent devices or integrated in one or more processors. A memory may also be provided in the processor 110 for storing instructions and data. For example, the processor 110 is used to implement the steps executed in the image processing method provided by the embodiments of the present application, and store instructions and data related to the image processing method.

[0082] The charging management module 140 is used to receive a charging input from a charger. Among them, the charger may be a wireless charger or a wired charger. The power management module 141 is used to connect the charging management module 140 and the processor 110.

[0083] The wireless communication function of the electronic device may be implemented by the antenna 1, the antenna 2, the mobile communication module 150, the wireless communication module 160, a modulation and demodulation processor, a baseband processor, etc.

[0084] The electronic device realizes the display function through the GPU, the display screen 194, the application processor, etc. The GPU is a microprocessor for image processing, connecting the display screen 194 and the application processor. The GPU is used to perform mathematical and geometric calculations for graphics rendering. For example, the GPU is used to execute the graphics rendering process in the image processing method.

[0085] The display screen 194 is used to display images, videos, etc. The display screen 194 includes a display panel. In some embodiments, the electronic device may include one or N display screens 194, where N is a positive integer greater than 1. For example, the display screen 194 is used to display the preview screen and the captured images in the camera application.

[0086] The electronic device can realize the shooting function through the ISP, the camera 193, the video codec, the GPU, the display screen 194, the application processor, etc.

[0087] The camera 193 is used to capture still images or videos. In some embodiments, the electronic device may include one or N cameras 193, where N is a positive integer greater than 1. For example, after the user's shooting operation, the camera 193 can be used to obtain the original image sequence.

[0088] The external memory interface 120 can be used to connect an external memory card, such as a Micro SD card, to expand the storage capacity of the electronic device. The internal memory 121 can be used to store computer-executable program code, and the executable program code includes instructions. The internal memory 121 may include a program storage area and a data storage area. For example, the internal memory 121 can be used to store the executable program code in the image processing method.

[0089] The electronic device can realize the audio function through the audio module 170, the speaker 170A, the receiver 170B, the microphone 170C, the headphone interface 170D, and the application processor, etc. For example, music playback, recording, etc.

[0090] The touch sensor can be disposed on the display screen 194, and the touch sensor and the display screen 194 form a touch screen, or "touch control screen". For example, the touch sensor is used to receive the trigger operation of the user for the shooting button.

[0091] The button 190 includes a power-on button, a volume button, etc. The button 190 can be a mechanical button or a touch button. The electronic device can receive button inputs and generate key signal inputs related to the user settings and function controls of the electronic device. In some scenarios, the electronic device can also respond to the user's operation on one or more buttons in the button 190 to implement the shooting operation, and the specific manner of the shooting operation in the embodiments of the present application is not limited.

[0092] The software system of an electronic device can adopt a layered architecture, an event-driven architecture, a microkernel architecture, a microservices architecture, or a cloud architecture, etc., which will not be elaborated here.

[0093] Figure 3 It is a schematic diagram of the software structure of the electronic device provided by the embodiments of the present application.

[0094] The layered architecture divides the system into several layers, and each layer has clear roles and divisions of labor. The layers communicate with each other through software interfaces. In some embodiments, the system is divided into five layers, from top to bottom are the application (APP), the application framework layer (FWK), the hardware abstraction layer (HAL), the driver layer, and the hardware layer, etc.

[0095] The application layer can include a series of application program packages. In the embodiments of the present application, the application program packages can include: camera, gallery, etc. The camera can implement image shooting and display of the taken pictures. The gallery, which can also be called an album, etc., can implement the storage and access of the taken pictures.

[0096] The application framework layer provides application programming interfaces (APIs) and programming frameworks for the application programs in the application layer. The application framework layer includes some predefined functions. In the embodiments of the present application, the application framework layer can include a camera access interface, where the camera access interface can include camera management and camera devices. The camera access interface is used to provide application programming interfaces and programming frameworks for camera applications.

[0097] The hardware abstraction layer is an interface layer between the application framework layer and the driver layer, providing a virtual hardware platform for the operating system. In the embodiments of the present application, the hardware abstraction layer can include a camera hardware abstraction layer and a camera algorithm library.

[0098] Among them, the camera hardware abstraction layer can provide virtual hardware for camera device 1, camera device 2, or more camera devices. The camera algorithm library can include the running code and data for implementing the image processing method provided by the embodiments of the present application.

[0099] The driver layer is the layer between hardware and software. The driver layer includes drivers for various hardware. The driver layer can include a camera device driver, a digital signal processor driver, and an image processor driver, etc.

[0100] Among them, the camera device driver is used to drive the sensor of the camera to collect images and drive the image signal processor to preprocess the images. The digital signal processor driver is used to drive the digital signal processor to process the images. The image processor driver is used to drive the graphics processor to process the images.

[0101] The following specifically describes the image processing method in the embodiments of the present application in combination with the above system structure:

[0102] In response to the user's operation of opening the camera application, such as clicking on the camera application icon, the camera application calls the camera access interface of the application framework layer to start the camera application, and then sends an instruction to start the camera by calling the camera device (such as a binocular camera) in the camera hardware abstraction layer. The camera hardware abstraction layer sends this instruction to the camera device driver in the kernel layer. This camera device driver can start the corresponding camera sensor and collect image optical signals through the sensor. One camera device in the camera hardware abstraction layer corresponds to one camera sensor in the hardware layer.

[0103] Then, the camera sensor can transmit the collected image optical signals to the image signal processor for preprocessing to obtain image electrical signals (original images), and transmit the above original images to the camera hardware abstraction layer through the camera device driver. The original images may include: a first image and a second image.

[0104] The camera hardware abstraction layer can send the first image and the second image to the camera algorithm library. The camera algorithm library stores program codes for implementing the image processing method provided in the embodiments of the present application. Based on the digital signal processor and the image processor, the camera algorithm library executes the above codes to implement the process of generating a target image based on the image processing of the first image and the second image.

[0105] The camera algorithm library can send the recognized original images collected by the camera to the camera hardware abstraction layer. Then, the camera hardware abstraction layer can display it.

[0106] It can be understood that the software architecture provided in the embodiments of the present application is only used as an example and does not constitute a limitation to the embodiments of the present application.

[0107] The following uses specific embodiments to elaborate in detail on the technical solutions of the present application and how the technical solutions of the present application solve the above technical problems. These several specific embodiments can be implemented independently or in combination with each other, and the same or similar concepts or processes may not be repeated in some embodiments.

[0108] Exemplarily, Figure 4 is a schematic flowchart of an image processing method provided in an embodiment of the present application. In Figure 4 the corresponding embodiment, an electronic device may include: a gallery, a camera access interface, a camera algorithm library, and a camera device driver. The functions of any module can be referred to in Figure 3 the corresponding embodiment and will not be elaborated here.

[0109] S401. In response to a photographing operation, the camera device driver acquires a first image and a second image.

[0110] The first image may be an image captured by a first camera, and the second image may be an image captured by a second camera.

[0111] A binocular camera may be provided in the electronic device. The binocular camera may include: a first camera and a second camera. When the first camera is the main camera in the binocular camera, the first image may be a left image; when the second camera is the auxiliary camera in the binocular camera, the second image may be a right image.

[0112] In response to a photographing operation, it includes: in response to a triggering operation of the user on the photographing button in the portrait photographing function, or in response to a triggering operation of the user on the photographing button in the HDR photographing function, or in response to a triggering operation of the user on the photographing button in the large aperture photographing function, or in response to a triggering operation of the user on the recording button in the video recording function, etc. In the embodiments of the present application, the applicable scenarios of the image processing method are not limited.

[0113] S402. The camera algorithm library acquires the first image and the second image from the camera device driver.

[0114] The camera hardware abstraction layer may acquire the first image and the second image from the camera device driver, and the camera algorithm library may acquire the first image and the second image from the camera hardware abstraction layer.

[0115] S403. The camera algorithm library uses the first image and the second image to calculate a first disparity map and a portrait segmentation map.

[0116] The camera algorithm library may calculate the first disparity map and the portrait segmentation map based on the following two methods (including Method 1 and Method 2).

[0117] Method 1: The camera algorithm library may perform feature matching on the first image and the second image to obtain a correlation map, and then calculate the first disparity map using the correlation map. The camera algorithm library may perform semantic segmentation on the first image and output a portrait segmentation map including the edges of the portrait.

[0118] Among them, the correlation map may be used to characterize the correlation between the first image and the second image. The correlation can be understood as the matching situation of feature points between two images. Among them, the feature matching method may include one or more of the following: based on speed up robust features (SURF), or scale-invariant feature transform (SIFT), or a neural network model, etc.

[0119] The human segmentation map can also be understood as a human mask map (or simply referred to as a mask map). In the mask map, the human region can be the region where the alpha value is equal to 255.

[0120] Method 2: A deep learning model can be pre-set in the camera algorithm library. The deep learning model is used to output a disparity map and a human segmentation map based on the input binocular images.

[0121] For example, the camera algorithm library can input the first image and the second image into the trained deep learning model and output the first disparity map and the human segmentation map. For specific details, please refer to Figure 6 the description in, which will not be elaborated here.

[0122] Among them, the deep learning model can perform semantic segmentation on the first image to calculate the human segmentation map corresponding to the first image. Moreover, the deep learning model can also determine the second disparity map based on the first image, the second image, and the human segmentation map.

[0123] S404. The camera algorithm library dilates the human segmentation map and determines the first edge region based on the human segmentation map and the dilated map.

[0124] The dilated map can be understood as the image obtained after performing dilation processing on the human segmentation map. The dilation processing method can include morphological filtering, etc.

[0125] The first edge region can be the image obtained by subtracting the dilated map from the human segmentation map by the camera algorithm library. The first edge region can include the human edges.

[0126] It can be understood that the electronic device can perform dilation processing so that the first edge region can include all the human edges in the human segmentation map. In this way, subsequently, the camera algorithm library can obtain the second edge region with the same position as the first edge region from the first disparity map and perform filtering processing on the second edge region to improve the accuracy of the human edges in the disparity map. Among them, the second edge region can be understood as the disparity region where the human edges need to be optimized.

[0127] Exemplarily, Figure 5 is an image schematic diagram provided by an embodiment of the present application. As Figure 5 shown, the human segmentation map can be the image shown in a of Figure 5 , the dilated map can be the image shown in b of Figure 5 , and the first edge region can be the black region shown in c of Figure 5 .

[0128] S405. The camera algorithm library uses the first edge region to obtain the second edge region from the first disparity map and performs filtering processing on the second edge region to obtain the third edge region.

[0129] For the meaning of the second edge region, reference may be made to the description in S404, which will not be elaborated here.

[0130] The filtering method may include one or more of the following: minimum filtering, Gaussian filtering, etc.

[0131] It can be understood that since the portrait segmentation map cannot guarantee edge fitting, the electronic device can obtain a more fitting edge region through filtering processing. In the disparity map, the closer an object is to the binocular camera, the greater its disparity. Therefore, by performing minimum filtering on a certain region 1 in the disparity map, the minimum value of the disparity map can be obtained, and this minimum value of the disparity map can be the disparity map of the background in region 1, so as to obtain a clearer and more fitting portrait edge. Gaussian filtering can be used to smooth the portrait edge, making the transition of the edge region relatively smooth.

[0132] In this way, through minimum filtering and Gaussian filtering of the second edge region, the portrait edge in the second edge region is closer to the portrait edge in the first edge region. After filtering processing, the portrait main body edge in the third edge region is sharper, so as to provide more accurate disparity information for the subsequent blurring effect.

[0133] S406. The camera algorithm library performs image fusion on the first disparity map and the third edge region to obtain a second disparity map.

[0134] For example, the camera algorithm library can perform image fusion on the region in the first disparity map except the third edge region and the third edge region to obtain a second disparity map with optimized portrait edges, improving the accuracy of the second disparity map.

[0135] S407. The camera algorithm library uses the second disparity map to perform blurring processing on the first image to obtain a target image.

[0136] The blurring processing can be methods such as Gaussian blur processing, and the embodiments of the present application do not make limitations in this regard.

[0137] S408. The gallery obtains the target image from the camera algorithm library.

[0138] For example, the camera algorithm library can return the processed target image to the gallery through modules such as a camera access interface.

[0139] S409. The gallery stores the target image.

[0140] It can be understood that Figure 4 the portrait segmentation map described in can not only be used to enhance the portrait edge in the disparity map, but also be applied to processing methods such as beauty skin processing, portrait enhancement processing, or matte extraction processing, etc. The embodiments of the present application do not make limitations on the usage scenarios of the portrait segmentation map.

[0141] Based on this, the electronic device can perform semantic segmentation, dilation processing, etc., to obtain a first edge region with a human portrait edge, and obtain a second edge region matching the first edge region from the first disparity map, and perform filtering processing on the second edge region to obtain a second disparity map with a clearer human portrait edge. After blurring the first image based on the second disparity map, a target image with accurate blurring of the human portrait edge can be obtained, reducing the situation of incorrect edge blurring or background leakage blurring caused by incorrect disparity estimation.

[0142] Based on the description in S403, the process of the camera algorithm library outputting the first disparity map and the human portrait segmentation map based on the deep learning model can be referred to the following Figure 6 description. Among them, the deep learning model can be a gated recurrent unit (GRU) iteration model or other models, etc.

[0143] Exemplarily, Figure 6 is a schematic diagram of a deep learning model provided by an embodiment of the present application. As Figure 6 shown, the deep learning model may involve one or more of the following: a left image feature encoding module 601, an ASPP module 602, a splicing module 603, a left image feature decoding module 604, a splicing module 605, a human portrait segmentation module 606, a right image feature encoding module 607, a correlation map calculation module 608, or a disparity estimation module 609.

[0144] The left image feature encoding module 601 is used to output a feature map corresponding to the first image, and the right image feature encoding module 607 is used to output a feature map corresponding to the second image. Among them, the left image feature encoding module 601 and the right image feature encoding module 607 can be feature encoding modules with the same structure, and the structure of the feature encoding module can be referred to Figure 7 the description in, which will not be elaborated here.

[0145] The ASPP module 602 is used for features, and the structure of the ASPP module can be referred to Figure 8 the description in, which will not be elaborated here.

[0146] Both the splicing module 603 and the splicing module 605 are used for image splicing. It can be understood that due to possible loss of some image information and increase of some image information after each image processing. Therefore, the higher the complexity of the image can be increased through image splicing, and the accuracy of subsequent processing based on the spliced image can be improved.

[0147] The left image feature decoding module 604 is used to decode the feature map and output a feature map with human portrait segmentation information. The structure of the left image feature decoding module (or called the feature decoding module) 604 can be referred to Figure 9The description in [reference] will not be repeated here.

[0148] The human portrait segmentation module 606 can be understood as a decoding module, that is, it realizes decoding the feature map into a human portrait segmentation map that can be recognized by the user. Among them, the structure of the human portrait segmentation module can be referred to Figure 10 The description in [reference] will not be repeated here.

[0149] Among them, the ASPP module 602, the splicing module 603, the left image feature decoding module 604, the splicing module 605, and the human portrait segmentation module 606 can jointly realize semantic segmentation of the image (the area shown by the dashed box).

[0150] The correlation map calculation module 608 is used to perform feature matching on the first image and the second image to obtain the matching relationship between feature points. In a deep learning model, the correlation map calculation module can generate a correlation feature map based on the input feature map, and this correlation feature map can be used to represent the feature matching relationship between feature maps.

[0151] The disparity estimation module 609 is used to calculate a disparity map based on the correlation feature map, the human portrait segmentation map, and the first image. Among them, the first image can be a reference image, so that the electronic device can be closer to the first image based on the first disparity map output by the disparity estimation module 609. The human portrait segmentation map contains human portrait segmentation information, and the human portrait segmentation information can represent the human portrait edge. The human portrait segmentation information can assist in the estimation of the first disparity map, so that the first disparity map can accurately identify the human portrait edge.

[0152] For example, when the correlation feature map indicates that feature 1 is determined to be the background, while the human portrait segmentation map indicates that feature 1 is a human portrait, the electronic device can take the human portrait segmentation map as the standard, determine that feature 1 is a human portrait, and then generate a first disparity map with accurate human portrait edges.

[0153] As Figure 6 shown, the process of obtaining the human portrait segmentation map corresponding to the first image can include the following steps. The camera algorithm library can input the first image into the left image feature encoding module 601, and encode to obtain the feature Figure 1 corresponding to the first image. The camera algorithm library inputs the feature Figure 1 into the ASPP module and outputs the feature Figure 2 . The camera algorithm library inputs the feature Figure 1 and the feature Figure 2 into the splicing module 603, and splices to obtain the feature Figure 3 . The camera algorithm library inputs the feature Figure 3 into the left image feature decoding module 604, and decodes to obtain the feature Figure 4 . The camera algorithm library inputs the feature Figure 1 and the feature Figure 4It is input into the splicing module 605 to splice and obtain features Figure 5 . The camera algorithm library inputs the features Figure 5 into the portrait segmentation module, and outputs a portrait segmentation map that can be recognized by the user. Among them, the features Figure 4 and the features Figure 5 can both contain portrait segmentation information.

[0154] The process of obtaining the first disparity map based on the first image and the second image may include the following steps. The camera algorithm library inputs the second image into the right image feature encoding module 607 to obtain features Figure 6 . The camera algorithm library inputs the features Figure 6 and the features Figure 1 into the correlation map calculation module 608, and outputs a correlation feature map. The camera algorithm library inputs the correlation feature map, the features Figure 1 , and the features Figure 5 into the disparity estimation module 609, and outputs the first disparity map.

[0155] It can be understood that since the deep learning model can implement the first disparity map based on the matching relationship between the first image and the second image, the portrait segmentation information, and the first image, the relatively accurate portrait edge can be recognized in the first disparity map.

[0156] Exemplarily Figure 7 is a schematic structural diagram of a feature encoding module provided by an embodiment of the present application.

[0157] As Figure 7 shown, the feature decoding module can be composed of 4 blocks, and each block can be composed of: convolution (conv) + instance normalization (IN) + activation function (relu), as shown by the dashed box. The stride in block 1 and block 3 can be 2, and the stride in block 2 and block 4 can be 1.

[0158] Among them, the camera algorithm library can also calculate the residual between the output result of block 1 and the output result of block 2, calculate the residual between the output result of block 2 and the output result of block 3, and the residual between the output result of block 3 and the output result of block 4, and perform better gradient descent through residual analysis.

[0159] Exemplarily, the camera algorithm library can sequentially input an image with a size of 1×3×H×W into 4 blocks, and output an image with a size of 1×C×H / 4×W / 4. The input image can be Figure 6 the first image or the second image described in Figure 7 . The camera algorithm library can obtain an output image with a total of C channels and 1 / 4 of the original size through 3 convolutions in .

[0160] It is understandable that Figure 7 the structure of the feature encoding module provided in [document] is only taken as an example and does not constitute a limitation to the embodiments of the present application.

[0161] Exemplarily, Figure 8 FIG. [X] is a schematic structural diagram of an ASPP module provided for an embodiment of the present application.

[0162] As Figure 8 shown, the ASPP module may be composed of multiple dilated convolutions with different sampling rates, such as a conv1×1 dilated convolution with a sampling rate (rate) of 0, a conv3×3 dilated convolution with a sampling rate of 1, a conv3×3 dilated convolution with a sampling rate of 5, and a conv3×3 dilated convolution with a sampling rate of 9. The above dilated convolutions with different sizes can be used to extract feature maps of different sizes. The ASPP module may further include steps such as pooling and upsampling.

[0163] Exemplarily, the camera algorithm library may input an image into multiple dilated convolutions respectively to obtain multiple output results. The multiple output results are concatenated and processed by a conv1×1 convolution to obtain an output image.

[0164] Generally, in order to fix the size of the input image, the electronic device may perform image processing such as cropping or stretching on the input image. However, the above processing methods result in a loss of image accuracy and affect the accuracy of subsequent semantic recognition results. In the embodiments of the present application, the ASPP module can extract image features at any size through multiple dilated convolutions with different sampling rates, and a more accurate semantic recognition result can be obtained based on the output result of the ASPP module.

[0165] It is understandable that Figure 8 the structure of the ASPP module provided in [document] is only taken as an example and does not constitute a limitation to the embodiments of the present application.

[0166] Exemplarily, Figure 9 FIG. [X] is a schematic structural diagram of a feature decoding module provided for an embodiment of the present application. The feature decoding module and the feature encoding module do not need to correspond one by one.

[0167] As Figure 9 shown, the feature decoding module may be composed of multiple blocks, such as a conv3×3 convolution block with a stride of 1, a conv1×1 convolution block with a stride of 1, and a bilinear upsampling module.

[0168] Exemplarily, the camera algorithm library may sequentially input an image with a size of 1×C×H / 4×W / 4 into the above-mentioned convolutional block and the bilinear upsampling module, and output an image with a size of 1×C×H / 2×W / 2.

[0169] It can be understood that Figure 9 the structure of the feature encoding module provided in [reference] is only taken as an example and does not constitute a limitation to the embodiments of the present application.

[0170] Exemplarily, Figure 10 is a schematic structural diagram of a portrait segmentation module provided for the embodiments of the present application.

[0171] As Figure 10 shown, the feature decoding module may be composed of multiple blocks, such as a convolutional block with conv3×3 and a group convolution value of C, a convolutional block with conv1×1 and a group convolution value of 1, a convolutional block with conv3×3 and a group convolution value of 1, and a bilinear upsampling module, etc.

[0172] Exemplarily, the camera algorithm library may sequentially input an image with a size of 1×C×H / 2×W / 2 into the above-mentioned convolutional block and the bilinear upsampling module, and output a portrait segmentation map with a size of 1×3×H×W.

[0173] It can be understood that Figure 10 the structure of the portrait segmentation module provided in [reference] is only taken as an example and does not constitute a limitation to the embodiments of the present application.

[0174] Combined with the description in the above Figures 4 - 10 in order to more clearly illustrate the embodiments of the present application, Figure 11 is a schematic diagram of the steps of an image processing method provided for the embodiments of the present application. As Figure 11 shown, the image processing method may include the following steps:

[0175] S1101. The electronic device acquires a portrait segmentation map and performs dilation on the portrait segmentation map to obtain a dilated map.

[0176] The portrait segmentation map may be determined based on the first image. The process by which the electronic device determines the portrait segmentation map based on the first image may refer to the description in S403 and will not be elaborated here.

[0177] S1102. The electronic device determines a portrait edge region based on the portrait segmentation map and the dilated map.

[0178] The portrait edge region may be Figure 4 the first edge region described in [reference]. The portrait edge region may be obtained by subtracting the portrait segmentation map from the dilated map.

[0179] S1103. The electronic device obtains a portrait edge area to be optimized that has the same position as the portrait edge area in the first parallax map.

[0180] The portrait edge area to be optimized can be Figure 4 the second edge area described in

[0181] S1104. The electronic device performs minimum filtering on the portrait edge area to be optimized.

[0182] S1105. The electronic device performs Gaussian filtering on the portrait edge area to be optimized.

[0183] Among them, for the effects of minimum filtering and Gaussian filtering, reference can be made to the description in S405, which will not be elaborated here.

[0184] After minimum filtering and Gaussian filtering processing, the electronic device can obtain an optimized portrait edge area. The optimized portrait edge area can be Figure 4 the third edge area described in

[0185] It should be noted that the module names involved in the embodiments of the present application can all be defined as other names, as long as the functions of each module can be realized, and no specific restrictions are imposed on the module names.

[0186] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in the embodiments of the present application are all information and data that have been authorized by the user or fully authorized by all parties. And the collection, use, and processing of relevant data need to comply with relevant laws, regulations, and standards of relevant countries and regions, and corresponding operation entrances are provided for users to choose to authorize or refuse.

[0187] The image processing method of the embodiments of the present application has been described above. Next, the device for executing the above method provided by the embodiments of the present application will be described. Those skilled in the art can understand that the method and the device can be combined and cited with each other. The relevant device provided by the embodiments of the present application can execute the steps in the method sorted above.

[0188] As Figure 12 shown, Figure 12 is a schematic structural diagram of an image processing device provided by an embodiment of the present application. This image processing device can be the electronic device in the embodiments of the present application, or a chip or chip system inside the electronic device.

[0189] AsFigure 12 As shown, the image processing device 1200 can be used in communication devices, circuits, hardware components, or chips. The image processing device 1200 includes: an acquisition unit 1201 and a processing unit 1202. Among them, the acquisition unit 1201 is used to support the step of data acquisition for executing the image processing method; the processing unit 1202 is used to support the image processing device 1200 to execute the information processing step.

[0190] In a possible implementation, the image processing device 1200 may further include a communication unit 1203. The communication unit 1203 is used to support the image processing device 1200 to execute steps such as receiving or sending messages.

[0191] The image processing devices described in the embodiments of this application may all include Figure 12 the units described in the corresponding embodiments.

[0192] Specifically, the processing unit 1202 and the acquisition unit 1201 may be integrated together, and communication may occur between the processing unit 1202 and the acquisition unit 1201.

[0193] In a possible implementation, the image processing device 1200 may further include: a storage unit 1204. Among them, the storage unit 1204 may include one or more memories. The memory may be a device or circuit in one or more devices used to store programs or data.

[0194] The storage unit 1204 may exist independently and be connected to the processing unit 1202 through a communication bus. The storage unit 1204 may also be integrated with the processing unit 1202.

[0195] Taking the image processing device 1200 as an example of the chip or chip system of the electronic device in the embodiments of this application, the storage unit 1204 may store computer-executable instructions of the method of the electronic device, so that the processing unit 1202 executes the method of the electronic device in the above embodiments. The storage unit 1204 may be a register, cache, or random access memory (RAM), etc. The storage unit 1204 may be integrated with the processing unit 1202. The storage unit 1204 may be a read-only memory (ROM) or other types of static storage devices that can store static information and instructions. The storage unit 1204 may be independent of the processing unit 1202.

[0196] In a possible implementation, the image processing device 1200 may further include: a communication unit 1203. The communication unit 1203 is used to support the interaction between the image processing device 1200 and other devices. Exemplarily, when the image processing device 1200 is an electronic device, the communication unit 1203 may be a communication interface or an interface circuit. When the image processing device 1200 is a chip or a chip system within an electronic device, the communication unit 1203 may be a communication interface. For example, the communication interface may be an input / output interface, a pin, or a circuit, etc.

[0197] The device in this embodiment can correspondingly be used to execute the steps performed in the above method embodiment, and its implementation principle and technical effects are similar, and will not be elaborated here.

[0198] Figure 13 This is a schematic diagram of the hardware structure of another electronic device provided by the embodiments of the present application.

[0199] The electronic device includes a processor 1301, a communication line 1304, and at least one communication interface ( Figure 13 exemplarily, the communication interface 1303 is taken as an example for illustration).

[0200] The processor 1301 may be a general-purpose central processing unit (CPU), a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits for controlling the execution of the program of the solution of the present application.

[0201] The communication line 1304 may include a circuit for transmitting information between the above components.

[0202] The communication interface 1303, using any device such as a transceiver, is used to communicate with other devices or communication networks, such as Ethernet, wireless local area networks (WLAN), etc.

[0203] Possibly, the electronic device may further include a memory 1302.

[0204] The memory 1302 can be a read-only memory (ROM) or other types of static storage devices that can store static information and instructions, a random access memory (RAM) or other types of dynamic storage devices that can store information and instructions, or can also be an electrically erasable programmable read-only memory (EEPROM), a compact disc read-only memory (CD-ROM), or other optical disc storage, optical disc storage (including compact discs, laser discs, optical discs, digital versatile discs, Blu-ray discs, etc.), magnetic disk storage media, or other magnetic storage devices, or any other medium that can be used to carry or store the desired program code in the form of instructions or data structures and can be accessed by a computer, but is not limited thereto. The memory can exist independently and be connected to the processor through the communication line 1304. The memory can also be integrated with the processor.

[0205] Among them, the memory 1302 is used to store computer execution instructions for executing the solution of this application, and is controlled by the processor 1301 to execute. The processor 1301 is used to execute the computer execution instructions stored in the memory 1302, so as to implement the method provided by the embodiments of this application.

[0206] Possibly, the computer execution instructions in the embodiments of this application can also be referred to as application program code, and the embodiments of this application do not make specific limitations thereto.

[0207] In a specific implementation, as an embodiment, the processor 1301 can include one or more CPUs, such as Figure 13 CPU0 and CPU1 in

[0208] In a specific implementation, as an embodiment, the electronic device can include multiple processors, such as Figure 13 the processor 1301 and the processor 1305 in

[0209] In the above embodiments, the instructions stored in the memory for the processor to execute can be implemented in the form of a computer program product. Among them, the computer program product can be pre-written in the memory, or can be downloaded and installed in the memory in the form of software.

[0210] The image processing method provided by the embodiments of the present application can be applied to an electronic device with communication functions. The electronic device includes a terminal device. For the specific device form of the terminal device, reference can be made to the above relevant description, which will not be elaborated here.

[0211] The embodiments of the present application provide a terminal device, which includes: a processor and a memory; the memory stores computer-executable instructions; the processor executes the computer-executable instructions stored in the memory, so that the terminal device executes the above method.

[0212] The embodiments of the present application provide a chip. The chip includes a processor, and the processor is used to call a computer program in the memory to execute the technical solutions in the above embodiments. Its implementation principle and technical effects are similar to those of the above relevant embodiments, which will not be elaborated here.

[0213] The embodiments of the present application also provide a computer-readable storage medium. The computer-readable storage medium stores a computer program. When the computer program is executed by a processor, the above method is implemented. The methods described in the above embodiments can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. If implemented in software, the functions can be stored as one or more instructions or codes on a computer-readable medium or transmitted on a computer-readable medium. The computer-readable medium can include a computer storage medium and a communication medium, and can also include any medium that can transmit a computer program from one place to another. The storage medium can be any target medium accessible by a computer.

[0214] In a possible implementation, the computer-readable medium can include RAM, ROM, a compact disc read-only memory (CD-ROM), or other optical disc memories, magnetic disk memories, or other magnetic storage devices, or any other medium targeted at carrying or storing the required program code in the form of instructions or data structures and accessible by a computer. Moreover, any connection is properly referred to as a computer-readable medium. For example, if software is transmitted from a website, server, or other remote source using coaxial cable, fiber optic cable, twisted pair, digital subscriber line (DSL), or wireless technologies such as infrared, radio, and microwave, then the coaxial cable, fiber optic cable, twisted pair, DSL, or wireless technologies such as infrared, radio, and microwave are included in the definition of the medium. As used herein, disk and optical disc include optical disc, laser disc, optical disc, digital versatile disc (DVD), floppy disk, and Blu-ray disc, where disks usually reproduce data magnetically, while optical discs use lasers to reproduce data optically. The above combinations should also be included within the scope of the computer-readable medium.

[0215] An embodiment of the present application provides a computer program product, which includes a computer program. When the computer program is run, it causes a computer to execute the above method.

[0216] The embodiments of the present application are described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or block in the flowchart and / or block diagram, and the combination of processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processing unit of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable devices to generate a machine, so that the instructions executed by the processing unit of the computer or other programmable data processing devices generate a device for implementing the functions specified in one process Figure 1 one process or multiple processes and / or blocks Figure 1 a block or multiple blocks.

[0217] The above specific implementation manners further elaborate on the purpose, technical solution, and beneficial effects of the present invention. It should be understood that the above are only specific implementation manners of the present invention and are not used to limit the protection scope of the present invention. Any modifications, equivalent replacements, improvements, etc. made on the basis of the technical solution of the present invention should be included in the protection scope of the present invention.

Claims

1. An image processing method, characterized in that, Including: In response to a photographing operation, obtain a first image and a second image, where the first image and the second image are images obtained based on a binocular camera; Obtain a first disparity map and an edge region in the first image, where the first disparity map is determined based on the visual difference between the first image and the second image, and the edge region is a region formed by the edges of the human figure in the first image; Obtain a first region in the first disparity map that has the same position as the edge region, and perform filtering processing on the first region to obtain a second region; Fuse the second region into the first disparity map to obtain a second disparity map; Use the second disparity map to perform blurring processing on the first image to obtain a target image.

2. The method according to claim 1, wherein Before obtaining the first disparity map and the edge region in the first image, the method further includes: Perform semantic segmentation on the first image to obtain a third image with a human figure region; Perform dilation processing on the third image to obtain a fourth image; Use the third image and the fourth image to determine the edge region.

3. The method according to claim 1 or 2, characterized in that, The performing filtering processing on the first region to obtain a second region includes: Perform minimum filtering and / or Gaussian filtering processing on the first region to obtain the second region.

4. The method according to any one of claims 1-3, characterized in that The obtaining the first disparity map includes: Input the first image and the second image into a deep learning model, and output the first disparity map. The deep learning model is used to determine the first disparity map by using the human figure segmentation information of the first image, the first image, and the matching relationship between the first image and the second image.

5. The method according to claim 4, wherein The deep learning model further includes: an Atrous Spatial Pyramid Pooling (ASPP) module, a feature encoding module, and a feature decoding module; the feature encoding module is used to perform feature extraction on the first image and the second image respectively to obtain a first feature map corresponding to the first image and a second feature map corresponding to the second image, the ASPP module is used to perform feature extraction on the first feature map to obtain a third feature map, the feature decoding module is used to decode the third feature map into a fourth feature map, the fourth feature map contains the human figure segmentation information, and the correlation calculation module is used to perform feature matching on the first feature map and the second feature map to obtain the matching relationship.

6. An image processing method, characterized in that, Including: Obtain a first image and a second image, where the first image and the second image are images obtained based on a binocular camera; Input the first image and the second image into a deep learning model, and output a first disparity map. The deep learning model is used to determine the first disparity map by using the human figure segmentation information of the first image, the first image, and the matching relationship between the first image and the second image.

7. The method according to claim 6, wherein The deep learning model includes: a correlation map calculation module and a disparity estimation module, The correlation map calculation module is used to perform feature matching on the first image and the second image to obtain the matching relationship, and the disparity estimation module is used to determine the first disparity map by using the human segmentation information and the matching relationship.

8. The method according to claim 7, wherein The deep learning model further includes: an atrous spatial pyramid pooling (ASPP) module, a feature encoding module, and a feature decoding module; The feature encoding module is used to respectively perform feature extraction on the first image and the second image to obtain a first feature map corresponding to the first image and a second feature map corresponding to the second image. The ASPP module is used to perform feature extraction on the first feature map to obtain a third feature map. The feature decoding module is used to decode the third feature map into a fourth feature map, and the fourth feature map contains the human segmentation information. The correlation calculation module is used to perform feature matching on the first feature map and the second feature map to obtain the matching relationship.

9. The method according to claim 8, wherein The deep learning model further includes: a first splicing module and a second splicing module. The method further includes: The first splicing module is used to splice the third feature map and the first feature map to obtain a fifth feature map; the second splicing module is used to splice the fourth feature map and the first feature map to obtain a sixth feature map, and the sixth feature map contains the human segmentation information.

10. An electronic device, characterized in that, The electronic device includes: one or more processors and a memory; The memory is coupled to the one or more processors. The memory is used to store computer program code, and the computer program code includes computer instructions. The one or more processors call the computer instructions to cause the electronic device to execute the method according to any one of claims 1 to 9.

11. A chip system, characterized in that, The chip system is applied to an electronic device. The chip system includes one or more processors, and the one or more processors are used to call computer instructions to cause the electronic device to execute the method according to any one of claims 1 to 9.

12. A computer-readable storage medium, characterized in that, The computer-readable storage medium includes computer instructions, and when the computer instructions run on an electronic device, the electronic device is caused to execute the method according to any one of claims 1 to 9.

13. A computer program product, characterized in that, The computer program product includes computer program code, and when the computer program code runs on an electronic device, the electronic device is caused to execute the method according to any one of claims 1 to 9.

Citation Information

Patent Citations

  • Stereo matching method and device

    CN110287964A

  • Image blurring method and device, storage medium and terminal equipment

    CN115170383A

  • Disparity map generation method and device, storage medium and terminal equipment

    CN115409759A

  • Method, apparatus and computer program product for disparity estimation

    EP2874395A2

  • Methods and systems using depth imaging for training and deploying neural networks for biometric Anti-spoofing

    WO2023065038A1