Image Processing Method and Apparatus

By using pixel classification maps, transparency maps and mask images for local feature enhancement and fusion processing in ultra-high resolution image processing, the problem of difficulty in efficiently estimating the transparency map of ultra-high resolution portrait images under limited memory and computing power in the prior art is solved, and a more efficient image cutting effect is achieved.

CN114299088BActive Publication Date: 2025-06-10BEIJING DAJIA INTERNET INFORMATION TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111611517.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-12-27
Publication Date
2025-06-10
Estimated Expiration
2041-12-27

AI Technical Summary

Technical Problem

The prior art is difficult to efficiently estimate the transparency map of ultra-high resolution portrait images under limited memory and computing power, resulting in errors or flaws in the prediction results.

Method used

By acquiring the pixel classification map, the first transparency map and the mask image of the input image, local feature enhancement is performed based on the mask image, the pixel classification map and the first transparency map, a second transparency map is generated, and a third transparency map is generated by fusing the first transparency map and the second transparency map, and finally segmenting the input image based on the third transparency map.

Benefits of technology

It improves the speed and effect of image processing, reduces the calculation amount of ultra-high resolution image cutouts, and improves the effect of cutouts.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114299088B_ABST
    Figure CN114299088B_ABST
Patent Text Reader

Abstract

The present disclosure relates to an image processing method and apparatus. The image processing method includes: obtaining an input image and a pixel classification map, a first transparency map, and a mask image corresponding to the input image; performing local feature enhancement on the input image based on the mask image, the pixel classification map, and the first transparency map to generate a second transparency map; performing a fusion process on the first transparency map and the second transparency map based on the mask image to generate a third transparency map of the input image; and segmenting the input image based on the third transparency map. According to the image processing method and apparatus of the present disclosure, the speed and effect of image processing can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of image processing technology. More specifically, the present disclosure relates to an image processing method and apparatus. Background Art

[0002] Portrait matting is a major technology in image editing and plays an important role in many applications such as photo editing, e-commerce demonstrations, and post-production of movies. The goal of portrait matting is to extract a fine transparency map of the portrait from a given image, and the foreground of the extracted portrait can be used to synthesize a new picture with other background images.

[0003] In addition to color pictures, traditional portrait matting algorithms usually rely on additional cues provided by users to mark the approximate area of the portrait, thereby reducing the uncertainty of matting and the complexity of transparency map estimation. However, due to the uneven quality of these external cues, the reliance on these external cues often leads to unstable prediction results or limited application scenarios. To solve this problem, recent portrait matting methods no longer rely on external cues, but instead determine the specific position of the portrait in the picture through portrait semantics or image saliency. These methods generally include two stages: the first step is to segment the portrait from the entire image; the second step is to consider local regions to refine the boundaries of the human body. By training on large-scale public or private portrait matting datasets, the above methods can obtain reasonable transparency map predictions without any user input.

[0004] However, relevant portrait matting methods mainly target low / medium-resolution images (usually not higher than 1080p). However, with the emergence of high-definition (HD) display devices, many applications such as game or movie production require the use of ultra-high-resolution (UHR) images. Given the limited memory and computing power of most portable devices and medium / low-end server devices, it is not easy to directly run previous portrait matting algorithms on ultra-high-resolution images. Although the previous matting algorithms can be simply extended, that is, the ultra-high-resolution image is divided into small regions, and forward inference is performed on each region based on the region, due to the lack of global information, the transparency map predicted by this scheme is prone to errors or defects. Therefore, how to efficiently estimate the transparency map of ultra-high-resolution portrait images under the conditions of limited memory and computing power has become an important issue to be solved urgently. Summary of the Invention

[0005] Exemplary embodiments of the present disclosure are provided to solve at least the problems in image processing in related technologies, or may not solve any of the above problems.

[0006] According to an exemplary embodiment of the present disclosure, there is provided an image processing method, including: obtaining an input image and a pixel classification map, a first transparency map, and a mask image corresponding to the input image, wherein the pixel classification map indicates the category of each pixel in the input image, and the category indicates whether each pixel is a transparent pixel or a non-transparent pixel, and the mask image indicates the area for image processing; performing local feature enhancement on the input image based on the mask image, the pixel classification map, and the first transparency map to generate a second transparency map; performing a fusion process on the first transparency map and the second transparency map based on the mask image to generate a third transparency map of the input image, wherein the resolution of the third transparency map is equal to the resolution of the input image; segmenting the input image based on the third transparency map.

[0007] Optionally, the step of performing local feature enhancement on the input image based on the mask image, the pixel classification map, and the first transparency map includes: respectively determining the areas to be optimized in the input image, the pixel classification map, and the first transparency map based on the mask image, to obtain the area to be optimized in the input image, the area to be optimized in the pixel classification map, and the area to be optimized in the first transparency map; fusing the area to be optimized in the input image, the area to be optimized in the pixel classification map, and the area to be optimized in the first transparency map to obtain global information; and performing local feature enhancement on the input image by inputting the global information into a preset local feature enhancement network.

[0008] Optionally, the step of performing a fusion process on the first transparency map and the second transparency map based on the mask image includes: determining the mask value of each pixel in the mask image as the weight of each pixel in the second transparency map; determining the difference between a preset value and the mask value of each pixel in the mask image as the weight of each pixel in the first transparency map; and performing weighted summation on the first transparency map and the second transparency map based on the weight of each pixel in the first transparency map and the weight of each pixel in the second transparency map.

[0009] Optionally, the step of obtaining the pixel classification map, the first transparency map, and the mask image corresponding to the input image includes: obtaining the pixel classification map of the input image and the first transparency map of the input image; and generating the mask image based on the first transparency map and the pixel classification map.

[0010] Optionally, the step of obtaining the pixel classification map of the input image and the first transparency map of the input image includes: obtaining a downsampled image corresponding to the input image; obtaining the pixel classification map of the input image and the first transparency map of the input image from the downsampled image, wherein the resolution of the first transparency map is smaller than the resolution of the input image.

[0011] Optionally, the step of obtaining the first transparency map of the input image from the downsampled image includes: extracting semantic features of at least one level of the downsampled image; predicting the position of the object to be segmented and the boundary of the object to be segmented in the input image based on the semantic features of at least one level; fusing the position of the object to be segmented and the boundary of the object to be segmented, and performing upsampling on the fusion result to obtain the first transparency map of the input image.

[0012] Optionally, the step of obtaining the pixel classification map of the input image from the downsampled image includes: predicting the category of each pixel in the input image based on the semantic features of at least one level; determining the pixel classification map of the downsampled image based on the category of each pixel in the input image; and performing upsampling on the pixel classification map of the downsampled image to obtain the pixel classification map of the input image.

[0013] Optionally, the step of obtaining the pixel classification map of the input image and the first transparency map of the input image from the downsampled image includes: obtaining the pixel classification map of the input image and the first transparency map of the input image from the downsampled image based on a preset global network.

[0014] Optionally, the global network includes an encoder and a decoder. The encoder is configured to extract semantic features of at least one level of the downsampled image, and the decoder is configured to perform feature reconstruction step by step using six levels, wherein the first three levels in the decoder are used to predict the position of the object to be segmented, and the last three levels in the decoder are used to predict the boundary of the object to be segmented.

[0015] Optionally, the local feature enhancement network includes a local enhancement layer, and the local enhancement layer is used to enhance local features based on global information.

[0016] Optionally, the local enhancement layer enhances local features by learning channel attention weights, wherein the channel attention weights are used to multiply with local features to encode foreground information and background information in global information into local features.

[0017] According to an exemplary embodiment of the present disclosure, there is provided an image processing apparatus, including: an acquisition unit configured to acquire an input image, a pixel classification map corresponding to the input image, a first transparency map, and a mask image, wherein the pixel classification map indicates the category of each pixel in the input image, the category indicating whether each pixel is a transparent pixel or a non-transparent pixel, and the mask image indicates a region for image processing; a second transparency map generation unit configured to perform local feature enhancement on the input image based on the mask image, the pixel classification map, and the first transparency map to generate a second transparency map; a third transparency map generation unit configured to perform a fusion process on the first transparency map and the second transparency map based on the mask image to generate a third transparency map of the input image, wherein the resolution of the third transparency map is equal to the resolution of the input image; and an image segmentation unit configured to segment the input image based on the third transparency map.

[0018] Optionally, the second transparency map generation unit is configured to: respectively determine regions to be optimized in the input image, the pixel classification map, and the first transparency map based on the mask image to obtain regions to be optimized in the input image, regions to be optimized in the pixel classification map, and regions to be optimized in the first transparency map; fuse the regions to be optimized in the input image, the regions to be optimized in the pixel classification map, and the regions to be optimized in the first transparency map to obtain global information; and perform local feature enhancement on the input image by inputting the global information into a preset local feature enhancement network.

[0019] Optionally, the third transparency map generation unit is configured to: determine the mask value of each pixel in the mask image as the weight of each pixel in the second transparency map; determine the difference between a preset value and the mask value of each pixel in the mask image as the weight of each pixel in the first transparency map; and perform a weighted sum of the first transparency map and the second transparency map based on the weight of each pixel in the first transparency map and the weight of each pixel in the second transparency map.

[0020] Optionally, the acquisition unit is configured to: acquire a pixel classification map of the input image and a first transparency map of the input image; and generate the mask image based on the first transparency map and the pixel classification map.

[0021] Optionally, the acquisition unit is configured to: acquire a downsampled image corresponding to the input image; and acquire the pixel classification map of the input image and the first transparency map of the input image from the downsampled image, wherein the resolution of the first transparency map is less than the resolution of the input image.

[0022] Optionally, the obtaining unit is configured to: extract semantic features of at least one level of the downsampled image; predict the position of the object to be segmented and the boundary of the object to be segmented in the input image based on the semantic features of at least one level; fuse the position of the object to be segmented and the boundary of the object to be segmented, and upsample the fusion result to obtain a first transparency map of the input image.

[0023] Optionally, the obtaining unit is configured to: predict the category of each pixel in the input image based on the semantic features of at least one level; determine a pixel classification map of the downsampled image based on the category of each pixel in the input image; and upsample the pixel classification map of the downsampled image to obtain a pixel classification map of the input image.

[0024] Optionally, the obtaining unit is configured to obtain a pixel classification map of the input image and a first transparency map of the input image from the downsampled image based on a preset global network.

[0025] Optionally, the global network includes an encoder and a decoder. The encoder is configured to extract semantic features of at least one level of the downsampled image. The decoder is configured to perform feature reconstruction step by step using six levels. Among them, the first three levels in the decoder are used to predict the position of the object to be segmented, and the last three levels in the decoder are used to predict the boundary of the object to be segmented.

[0026] Optionally, the local feature enhancement network includes a local enhancement layer, and the local enhancement layer is used to enhance local features based on global information.

[0027] Optionally, the local enhancement layer enhances local features by learning channel attention weights, where the channel attention weights are used to multiply with local features to encode foreground information and background information in global information into local features.

[0028] According to an exemplary embodiment of the present disclosure, there is provided an electronic device, including: a processor; a memory for storing executable instructions of the processor; wherein, the processor is configured to execute the instructions to implement an image processing method according to an exemplary embodiment of the present disclosure.

[0029] According to an exemplary embodiment of the present disclosure, there is provided a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor of an electronic device, the electronic device is enabled to execute an image processing method according to an exemplary embodiment of the present disclosure.

[0030] According to an exemplary embodiment of the present disclosure, there is provided a computer program product including a computer program / instructions, which, when executed by a processor, implement an image processing method according to an exemplary embodiment of the present disclosure.

[0031] The technical solutions provided by the embodiments of the present disclosure at least bring the following beneficial effects:

[0032] Improve the speed and effect of image processing;

[0033] For ultra-high resolution images, reduce the computational amount of image matting and improve the effect of image matting.

[0034] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS

[0035] The accompanying drawings herein are incorporated into the specification and constitute a part of this specification, showing embodiments consistent with the present disclosure, and together with the specification are used to explain the principles of the present disclosure and do not constitute an improper limitation of the present disclosure.

[0036] Figure 1 An exemplary system architecture in which the exemplary embodiments of the present disclosure can be applied is shown.

[0037] Figure 2 A flowchart showing an image processing method according to an exemplary embodiment of the present disclosure is shown.

[0038] Figure 3 A structural diagram showing an image processing framework according to an exemplary embodiment of the present disclosure is shown.

[0039] Figure 4A A simplified structural diagram showing a local feature enhancement network in an image processing framework according to an exemplary embodiment of the present disclosure is shown.

[0040] Figure 4B An example of a masked shrinkage and dilation layer in a local feature enhancement network according to an exemplary embodiment of the present disclosure is shown.

[0041] Figure 5 A block diagram showing an image processing apparatus according to an exemplary embodiment of the present disclosure is shown.

[0042] Figure 6 It is a block diagram of an electronic device 600 according to an exemplary embodiment of the present disclosure. DETAILED DESCRIPTION

[0043] In order to enable those of ordinary skill in the art to better understand the technical solutions of the present disclosure, the technical solutions in the embodiments of the present disclosure will be clearly and completely described below with reference to the accompanying drawings.

[0044] It should be noted that the terms "first", "second", etc. in the description, claims and above-mentioned drawings of the present disclosure are used to distinguish similar objects, and do not necessarily have to be used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances, so that the embodiments of the present disclosure described here can be implemented in an order other than those illustrated or described here. The embodiments described in the following examples do not represent all embodiments consistent with the present disclosure. On the contrary, they are merely examples of devices and methods consistent with some aspects of the present disclosure as detailed in the appended claims.

[0045] It should be noted here that "at least one of several items" in the present disclosure all represents three parallel situations, including "any one of the several items", "a combination of any multiple of the several items", and "all of the several items". For example, "including at least one of A and B" includes the following three parallel situations: (1) including A; (2) including B; (3) including A and B. Another example is "performing at least one of step one and step two", which means the following three parallel situations: (1) performing step one; (2) performing step two; (3) performing step one and step two.

[0046] An ultra-high resolution image refers to an image with a resolution of 4K or higher, which far exceeds the resolution that existing matting algorithms can handle on portable devices.

[0047] In a related technology, first, the ultra-high resolution picture is downsampled, and then the downsampled picture. In order to achieve full-image forward inference, that is, input the entire picture into the matting network for prediction, the picture has to be downsampled to a size that the memory and computing power of the computing device can accept. However, downsampling the picture will cause information loss. Especially in the fine-structured hair area, downsampling will cause the hair strands to blur, making it difficult to extract an accurate transparency map.

[0048] In another related technology, a transparency map of an ultra-high resolution image including a portrait picture is obtained by using region-based forward inference. Specifically, the given ultra-high resolution image is cut into multiple small regions, and the existing portrait matting algorithm is run on each region separately. Finally, the transparency prediction results of each piece are combined to obtain the transparency map of the final full image. However, this method has at least two defects. First, due to the inability to communicate information between each region and the lack of global information, the transparencies predicted for different regions are not exactly the same, and obvious region boundaries can even be seen in the transparency map of the final combined full image. Second, in order to improve the consistency of predictions between regions as much as possible, there will be an overlap between adjacent regions when dividing the regions. However, this approach will increase the computational complexity and time consumption of forward inference for a single picture, which is not advisable.

[0049] The present disclosure aims to explore how to efficiently obtain a high-precision transparency map of a portrait image under the condition of limited memory and computing power of a computing device. Most regions in a portrait image belong to the determined background region and foreground region, and only a very small number of regions belong to the transition region, which is also the most concerned region in image matting. Therefore, by removing large chunks of the background region and foreground region and only retaining the most concerned transition region, the amount of computation can be significantly reduced, and sparse convolution can be used to achieve full-image inference under the condition of limited memory.

[0050] Next, a description will be made with reference to Figures 1 to 6 specifically describe an image processing method and apparatus according to an exemplary embodiment of the present disclosure.

[0051] Figure 1 An exemplary system architecture 100 in which an exemplary embodiment of the present disclosure can be applied is shown.

[0052] As Figure 1 shown, the system architecture 100 may include terminal devices 101, 102, 103, a network 104, and a server 105. The network 104 is used to provide a medium for communication links between the terminal devices 101, 102, 103 and the server 105. The network 104 may include various connection types, such as wired, wireless communication links, or fiber optic cables, etc. Users can use the terminal devices 101, 102, 103 to interact with the server 105 through the network 104 to receive or send messages (such as video data upload requests, video data download requests), etc. Various communication client applications may be installed on the terminal devices 101, 102, 103, such as audio and video call software, audio and video recording software, instant communication software, conference software, email clients, social platform software, etc. The terminal devices 101, 102, 103 may be hardware or software. When the terminal devices 101, 102, 103 are hardware, they may be various electronic devices with a display screen and capable of audio and video playback, recording, editing, etc., including but not limited to smart phones, tablet computers, laptop portable computers, and desktop computers, etc. When the terminal devices 101, 102, 103 are software, they may be installed in the above-listed electronic devices, and may be implemented as multiple software or software modules (for example, used to provide distributed services), or may be implemented as a single software or software module. No specific limitation is made here.

[0053] The terminal devices 101, 102, and 103 may be installed with image acquisition devices (e.g., cameras) to acquire video data. In practice, the smallest visual unit that makes up a video is a frame. Each frame is a static image. Synthesizing a sequence of temporally continuous frames together forms a dynamic video. In addition, the terminal devices 101, 102, and 103 may also be installed with components (e.g., speakers) for converting electrical signals into sounds to play sounds, and may also be installed with devices (e.g., microphones) for converting analog audio signals into digital audio signals to acquire sounds.

[0054] The server 105 may be a server that provides various services, such as a background server that supports multimedia applications installed on the terminal devices 101, 102, and 103. The background server may parse, store, and process data such as audio-visual data upload requests received, and may also receive audio-visual data download requests sent by the terminal devices 101, 102, and 103, and feedback the audio-visual data indicated by the audio-visual data download requests to the terminal devices 101, 102, and 103.

[0055] It should be noted that the server may be hardware or software. When the server is hardware, it may be implemented as a distributed server cluster composed of multiple servers, or may be implemented as a single server. When the server is software, it may be implemented as multiple software or software modules (e.g., for providing distributed services), or may be implemented as a single software or software module. No specific limitation is made here.

[0056] It should be noted that the image processing method provided by the embodiments of the present disclosure is usually executed by the terminal device, but may also be executed by the server, or may be executed by the cooperation of the terminal device and the server. Correspondingly, the image processing device may be set in the terminal device, in the server, or in both the terminal device and the server.

[0057] It should be understood that Figure 1 the numbers of the terminal devices, networks, and servers in

[0058] Figure 2 FIG. shows a flowchart of an image processing method according to an exemplary embodiment of the present disclosure.

[0059] Referring to Figure 2 , in step S201, an input image and a pixel classification map, a first transparency map, and a mask image corresponding to the input image are acquired. Here, the pixel classification map indicates the category of each pixel in the input image, and the category indicates whether each pixel is a transparent pixel or a non-transparent pixel, and the mask image indicates the area for performing image processing.

[0060] In an exemplary embodiment of the present disclosure, when obtaining the pixel classification map, the first transparency map, and the mask image corresponding to the input image, the pixel classification map of the input image and the first transparency map of the input image may be obtained first, and then the mask image may be generated based on the first transparency map and the pixel classification map, so as to obtain a mask image for removing large background regions and foreground regions of the image and only retaining the most concerned transition regions. Here, the pixel classification map represents the category of each pixel in the picture, distinguishing whether a pixel is a transparent pixel (i.e., the transparency value is between 0 and 1) or a non-transparent pixel (i.e., the transparency value is 0 or 1). Here, the mask image may be obtained by performing dilation and erosion operations on this rough transparency, and the unknown regions in the mask image define the regions that need to be re-optimized.

[0061] In an exemplary embodiment of the present disclosure, when obtaining the pixel classification map of the input image and the first transparency map of the input image, the downsampled image corresponding to the input image may be obtained first, and then the pixel classification map of the input image and the first transparency map of the input image may be obtained from the downsampled image, so as to reduce the computational cost. Here, the resolution of the first transparency map is less than the resolution of the input image.

[0062] In an exemplary embodiment of the present disclosure, when obtaining the first transparency map of the input image from the downsampled image, at least one level of semantic features of the downsampled image may be extracted first, and then the positions and boundaries of the objects to be segmented in the input image may be predicted based on the at least one level of semantic features. Finally, the positions and boundaries of the objects to be segmented are fused, and the fusion result is upsampled to obtain the first transparency map of the input image, so as to obtain the first transparency map based on the semantic features.

[0063] In an exemplary embodiment of the present disclosure, when obtaining the pixel classification map of the input image from the downsampled image, the category of each pixel in the input image may be predicted first based on the at least one level of semantic features, and then the pixel classification map of the downsampled image may be determined based on the category of each pixel in the input image. The pixel classification map of the downsampled image is upsampled to obtain the pixel classification map of the input image, so as to determine the pixel classification map based on the category of each pixel in the input image.

[0064] In an exemplary embodiment of the present disclosure, when obtaining the pixel classification map of the input image and the first transparency map of the input image from the downsampled image, the pixel classification map of the input image and the first transparency map of the input image may be obtained from the downsampled image based on a preset global network, so as to obtain the pixel classification map and the first transparency map through the global network. That is, the preset global network is a multi-scale multi-task global network.

[0065] In an exemplary embodiment of the present disclosure, the global network may include an encoder and a decoder. The encoder is configured to extract semantic features of at least one level of the downsampled image, and the decoder is configured to perform feature reconstruction step by step using six levels. Here, the first three levels in the decoder are used to predict the position of the object to be segmented, and the last three levels in the decoder are used to predict the boundary of the object to be segmented. The above configuration of the global network enables the global network to be used to obtain a pixel classification map and a first transparency map.

[0066] In step S202, based on the mask image, the pixel classification map, and the first transparency map, local feature enhancement is performed on the input image to generate a second transparency map. Here, the resolution of the second transparency map may be greater than or equal to the resolution of the input image.

[0067] In an exemplary embodiment of the present disclosure, when performing local feature enhancement on the input image based on the mask image, the pixel classification map, and the first transparency map, the regions to be optimized in the input image, the pixel classification map, and the first transparency map may be determined respectively based on the mask image to obtain the regions to be optimized in the input image, the regions to be optimized in the pixel classification map, and the regions to be optimized in the first transparency map, and the regions to be optimized in the input image, the regions to be optimized in the pixel classification map, and the regions to be optimized in the first transparency map are fused to obtain global information, and then the global information is input into a preset local feature enhancement network to perform local feature enhancement on the input image, thereby realizing local feature enhancement of the input image. Here, the unknown regions in the mask image can be used to index the input image, the pixel classification map, and the first transparency map (that is, the foreground and background regions are discarded, and only the unknown regions are retained), and sparse representations of the input image, the pixel classification map, and the first transparency map in the unknown region part are obtained.

[0068] In an exemplary embodiment of the present disclosure, the local feature enhancement network may include a local enhancement layer, and the local enhancement layer may be used to enhance local features based on global information. Here, the local feature enhancement network may be, for example, a local network based on sparse convolution. The above configuration of the local feature enhancement network enables the local feature enhancement network to be used to perform local feature enhancement on the input image.

[0069] In an exemplary embodiment of the present disclosure, the local enhancement layer may enhance local features by learning channel attention weights. Here, the channel attention weights may be used to multiply with the local features to encode the foreground information and background information in the global information into the local features. The above function of the local enhancement layer enables the local feature enhancement network to be used to perform local feature enhancement on the input image.

[0070] In step S203, the first transparency map and the second transparency map are fused based on the mask image to generate a third transparency map of the input image. Here, the resolution of the third transparency map is equal to the resolution of the input image.

[0071] In an exemplary embodiment of the present disclosure, when fusing the first transparency map and the second transparency map based on the mask image, the mask value of each pixel in the mask image can be first determined as the weight of each pixel in the second transparency map, and the difference between the preset value and the mask value of each pixel in the mask image can be determined as the weight of each pixel in the first transparency map. Then, the first transparency map and the second transparency map are weighted and summed based on the weight of each pixel in the first transparency map and the weight of each pixel in the second transparency map, so as to obtain a transparency map for image segmentation through weighted summation.

[0072] In step S204, the input image is segmented based on the third transparency map.

[0073] Figure 3 The structural diagram of an image processing framework according to an exemplary embodiment of the present disclosure is shown. Figure 3 The image processing framework in can be applicable to ultra-high resolution pictures. In Figure 3 the image processing framework is an efficient human matting network (EHM-Net for short). As Figure 3 shown, the efficient human matting network includes a multi-scale multi-task global network (Global Network, G-Net for short) and a local feature enhancement network.

[0074] The input of the multi-scale multi-task global network is the downsampled image of the input image. The multi-scale multi-task global network aims to extract a rough portrait transparency map from the input low-resolution image. The given ultra-high resolution image is first downsampled to, for example, but not limited to, 512×512. The multi-scale multi-task global network uses a base network (such as, for example, a lightweight deep neural network (Mobile Network, abbreviated as MobileNet), a residual network (Residual Network, abbreviated as ResNet), etc.) to extract semantic features at different levels. After obtaining these features, although the transparency map can be directly regressed through a decoder with skip connections, this is not the optimal strategy. In the absence of prior knowledge, portrait matting can be divided into two sub-tasks: pixel-level portrait localization (i.e., human body segmentation) and estimating the transparency value of the boundary (i.e., portrait matting). However, the goals of these two tasks are not exactly the same. Specifically, human body segmentation requires more global information, so a large receptive field is used to distinguish the pixels in the portrait area from the background. On the other hand, portrait matting pays more attention to local context and uses neighboring information to separate the foreground and background colors.

[0075] As Figure 3 shown, the decoder in the multi-scale multi-task global network uses six levels to perform feature reconstruction step by step. Different from traditional methods, in the decoder of the multi-scale multi-task global network, the first three low-resolution levels are used to predict the multi-scale human body segmentation results S i , i ∈ {1, 2, 3}, and the last three high-resolution levels are used to predict the matting results M i , i ∈ {4, 5, 6}. Then, the final result (transparency map) is obtained by fusing the human body segmentation result S 3 and the boundary matting result M 6 . When supervising human body segmentation, the mean squared error loss (L2 loss) can be used and the whole image is supervised; when supervising portrait matting, the mean absolute error loss (L1 loss) can be used but only the boundary area is supervised; when supervising the final result , the L1 loss and the Laplacian loss can be used.

[0076] In the decoder of the multi-scale multi-task global network, a pixel classification map (marked as SPC in Figure 3 ) can also be predicted, as shown in the upper right corner of Figure 3 . While the multi-scale multi-task global network predicts portrait segmentation and transparency, each layer of the decoder is connected to a fully connected layer to predict the multi-scale pixel classification map. When training the multi-scale multi-task global network, this prediction is supervised with a classification loss function. By introducing this classification task, the performance of the multi-scale multi-task global network can be improved.

[0077] After extracting the rough transparency map from the multi-scale multi-task global network, the local feature enhancement network can be used to progressively refine the rough transparency map, thereby reconstructing the original super high-resolution transparency map.

[0078] In super high-resolution images, most of the observed image regions consist of well-defined background or foreground pixels (pixels with transparency 0 or 1). According to the statistics of the portrait matting training set, the proportions of well-defined foreground pixels and transparent pixels are 38.6% and 3.8% respectively, indicating that a large amount of computation may be wasted on large foreground and background regions. To address this issue, a simple approach is to downsample the input image to a smaller scale to meet the memory limit. However, the information loss caused by downsampling leads to very blurred predictions by the model in the boundary regions, especially in the hair region. Another solution is region-based forward inference, that is, dividing the full image into several small regions, and only considering the regions containing unknown regions during forward inference while discarding other regions. However, when the divided regions are too small, due to the lack of context information of adjacent regions, the predictions between regions cannot be fully consistent; while when the divided regions are too large, a large amount of computational overhead will be introduced. Another solution is the replacement strategy based on small regions, that is, in the second stage, select small regions with lower confidence or higher error rate to be refined and then replace the corresponding regions in the rough transparency map. However, when the input image is super high-resolution, the resolution gap will result in obvious replacement defects in the prediction results.

[0079] Since the matting model can obtain sufficient information from neighboring regions rather than the full image, it is possible to skip well-defined foreground or background regions and only focus on local unknown regions, which will save a large amount of memory and computational volume. Different from the solutions mentioned above, in the present disclosure, a sparse convolutional network is used to reduce the computational volume while ensuring sufficient local information.

[0080] The transparency maps of portraits have a high degree of consistency in data distribution, that is, the transparency maps of portraits are composed of large trunk regions and narrow boundary regions. Generally, the transparency of the trunk region is 1, and the matting of the trunk is similar to human body segmentation, and does not require high-resolution image input. A resolution of 512p is sufficient to accurately segment the trunk part of the human body. The transparency of the boundary region is between 0 and 1. The boundary region can be a hard boundary (for example, the boundary of clothes, the boundary of the arm, etc.), or a soft boundary (for example, the boundary of hair). The soft boundary composed of fine structures usually has a high requirement for the resolution of the input picture. Therefore, the prediction of the original resolution transparency can be expressed as the following equation. In the following equation, is a binary mask, 0 indicates that the pixel is not considered, and are upsampling and downsampling operations, represents a low-resolution transparency map, i.e., a rough transparency map (e.g., the first transparency map), represents a high-resolution transparency map, i.e., a refined transparency map (e.g., the second transparency map), represents the final transparency map (e.g., the third transparency map) MaxPool k represents a max pooling operation with a kernel size of k. The subscripts h and l denote high resolution and low resolution, Θ are network parameters, the superscript G represents the global network for multi-scale multi-task, and the superscript L represents the local feature enhancement network.

[0081]

[0082]

[0083]

[0084] Here, can be estimated from the low-resolution input, while can be extracted from the high-resolution input. For is obtained from the global network for multi-scale multi-task. As for due to memory limitations, full-image inference is usually not feasible, so a local network based on sparse convolution is constructed, only focusing on the boundary regions.

[0085] For perform a "dilation" operation to generate a mask (i.e., a mask image) to be input into the local feature enhancement network to identify the pixels to be skipped. Since the foreground and background information in the neighboring regions is crucial for separating the foreground color from the background, it is not feasible to directly select the transparency pixels in as the sparse index. Here, max pooling with a kernel size of k can be used to simulate dilation to improve the computational efficiency on the graphics processor. If k is too small, it will result in insufficient information in the neighboring regions, while too large a k will introduce computational overhead. To balance accuracy and efficiency, k can be set to 25.

[0086] Figure 4A shows a simplified structural diagram of the local feature enhancement network in the image processing framework according to an exemplary embodiment of the present disclosure. Figure 4B shows an example of a masked Squeeze-and-Excitation (masked-SE) layer in the local feature enhancement network according to an exemplary embodiment of the present disclosure. As shown in FIG. 4, the local feature enhancement network includes a masked Squeeze-and-Excitation layer.

[0087] The local feature enhancement network can be a matte extraction network similar to a U - shape Network (U - Net for short), implemented through sparse convolution. For example, but not limited to, a sparse ResNet50 can be used as the backbone to extract multi - level encoded semantic features and used as skip layers in the sparse decoder. Before reconstructing the features through the decoder, global information can be extracted from the multi - scale multi - task global network and the local features can be enhanced through the above - mentioned masked contraction and dilation layers. As Figure 4B shown, the masked contraction and dilation layer learns channel attention weights, which are used to multiply with the local features to encode global foreground and background information into the local features. In Figure 4B where, represents the intermediate layer of the multi - scale multi - task global network, represents the low - resolution transparency map, GAP represents the global pooling layer, FC represents the fully - connected layer, represents the intermediate layer of the local feature enhancement network.

[0088] The input of the local feature enhancement network is the part corresponding to the unknown regions in the input image, the pixel classification map, and the rough transparency map (i.e., the first transparency map). Since the part corresponding to the unknown regions in the input image, the pixel classification map, and the rough transparency map (i.e., the first transparency map) is obtained by indexing the input image, the pixel classification map, and the rough transparency map (i.e., the first transparency map) with the unknown regions in the masked image (the masked image is labeled as g in Figure 3 ), therefore, the input dimension of the local feature enhancement network is N×C×K, rather than N×C×H×W. Here, N represents the number of input samples, C, H, and W represent the number of channels, length, and width of each sample respectively, and K represents the number of pixels to be processed.

[0089] In the exemplary embodiments of the present disclosure, the local feature enhancement network can be a sparse version of a common encoder - decoder structure. The difference between the local feature enhancement network and the common encoder - decoder network is that the convolutional layers in the common encoder - decoder network are implemented with dense convolution operations (i.e., the input dimension is N×C×H×W), but the convolutional layers in the local feature enhancement network are implemented with sparse convolution operations (i.e., the input dimension is N×C×K). After the encoder extracts features and the decoder reconstructs the features, A h .

[0090] Since no special design is made for the network structures of the auto - encoders in the global network of multi - scale multi - tasks and the local feature enhancement network in the exemplary embodiments of the present disclosure, different self - encoding - decoder networks can be switched according to different application scenarios. For example, a deeper network with higher accuracy can be used when applying on the server side, and a real - time lightweight network can be used when applying on the mobile side.

[0091] The above has been combined with Figures 1 to 4B to describe the image processing method according to the exemplary embodiments of the present disclosure. In the following, the image processing apparatus and its units according to the exemplary embodiments of the present disclosure will be described with reference to Figure 5 the image processing apparatus and its units according to the exemplary embodiments of the present disclosure will be described.

[0092] Figure 5 The block diagram showing the image processing apparatus according to the exemplary embodiments of the present disclosure.

[0093] With reference to Figure 5 , the image processing apparatus includes an acquisition unit 51, a second transparency map generation unit 52, a third transparency map generation unit 53, and an image segmentation unit 54.

[0094] The acquisition unit 51 is configured to acquire an input image, a pixel classification map corresponding to the input image, a first transparency map, and a mask image. Here, the pixel classification map indicates the category of each pixel in the input image, where the category indicates whether each pixel is a transparent pixel or a non - transparent pixel, and the mask image indicates the area for image processing.

[0095] In the exemplary embodiments of the present disclosure, the acquisition unit 51 may be configured to: acquire the pixel classification map of the input image and the first transparency map of the input image; generate a mask image based on the first transparency map and the pixel classification map.

[0096] In the exemplary embodiments of the present disclosure, the acquisition unit 51 may be configured to: acquire a down - sampled image corresponding to the input image; acquire the pixel classification map of the input image and the first transparency map of the input image from the down - sampled image. Here, the resolution of the first transparency map is less than that of the input image.

[0097] In the exemplary embodiments of the present disclosure, the acquisition unit 51 may be configured to: extract semantic features of at least one level of the down - sampled image; predict the position and boundary of the object to be segmented in the input image based on the semantic features of the at least one level; fuse the position and boundary of the object to be segmented, and upsample the fusion result to obtain the first transparency map of the input image.

[0098] In an exemplary embodiment of the present disclosure, the acquisition unit 51 may be configured to: predict the category of each pixel in the input image based on the semantic features of at least one level; determine a pixel classification map of the downsampled image based on the category of each pixel in the input image; and upsample the pixel classification map of the downsampled image to obtain a pixel classification map of the input image.

[0099] In an exemplary embodiment of the present disclosure, the acquisition unit 51 may be configured to: obtain a pixel classification map of the input image and a first transparency map of the input image from the downsampled image based on a preset global network.

[0100] In an exemplary embodiment of the present disclosure, the global network may include an encoder and a decoder. The encoder may be configured to extract semantic features of at least one level of the downsampled image, and the decoder may be configured to perform feature reconstruction step by step using six levels. Here, the first three levels in the decoder may be used to predict the position of the object to be segmented, and the last three levels in the decoder may be used to predict the boundary of the object to be segmented.

[0101] The second transparency map generation unit 52 is configured to perform local feature enhancement on the input image based on the mask image, the pixel classification map, and the first transparency map to generate a second transparency map.

[0102] In an exemplary embodiment of the present disclosure, the second transparency map generation unit 52 may be configured to: respectively determine the regions to be optimized in the input image, the pixel classification map, and the first transparency map based on the mask image to obtain the regions to be optimized in the input image, the regions to be optimized in the pixel classification map, and the regions to be optimized in the first transparency map; fuse the regions to be optimized in the input image, the regions to be optimized in the pixel classification map, and the regions to be optimized in the first transparency map to obtain global information; and perform local feature enhancement on the input image by inputting the global information into a preset local feature enhancement network. Here, the resolution of the second transparency map may be greater than or equal to the resolution of the input image.

[0103] In an exemplary embodiment of the present disclosure, the local feature enhancement network may include a local enhancement layer, and the local enhancement layer may be used to enhance local features based on global information.

[0104] In an exemplary embodiment of the present disclosure, the local enhancement layer may enhance local features by learning channel attention weights. Here, the channel attention weights may be used to multiply with local features to encode foreground information and background information in the global information into local features.

[0105] The third transparency map generation unit 53 is configured to perform a fusion process on the first transparency map and the second transparency map based on the mask image to generate a third transparency map of the input image. Here, the resolution of the third transparency map is equal to the resolution of the input image.

[0106] In an exemplary embodiment of the present disclosure, the third transparency map generation unit 53 may be configured to: determine the mask value of each pixel in the mask image as the weight of each pixel in the second transparency map; determine the difference between a preset value and the mask value of each pixel in the mask image as the weight of each pixel in the first transparency map; perform weighted summation on the first transparency map and the second transparency map based on the weight of each pixel in the first transparency map and the weight of each pixel in the second transparency map.

[0107] The image segmentation unit 54 is configured to segment the input image based on the third transparency map.

[0108] Regarding the device in the above embodiments, the specific manner in which each unit performs operations has been described in detail in the embodiments related to the method, and will not be elaborated here.

[0109] The above has been combined with Figure 5 to describe the image processing device according to the exemplary embodiments of the present disclosure. Next, in combination with Figure 6 to describe the electronic device according to the exemplary embodiments of the present disclosure.

[0110] Figure 6 is a block diagram of an electronic device 600 according to the exemplary embodiments of the present disclosure.

[0111] Referring to Figure 6 , the electronic device 600 includes at least one memory 601 and at least one processor 602. A set of computer-executable instructions is stored in the at least one memory 601. When the set of computer-executable instructions is executed by the at least one processor 602, the method for image processing according to the exemplary embodiments of the present disclosure is executed.

[0112] In an exemplary embodiment of the present disclosure, the electronic device 600 may be a PC computer, a tablet device, a personal digital assistant, a smart phone, or other devices capable of executing the above set of instructions. Here, the electronic device 600 does not have to be a single electronic device, and may also be any assembly of devices or circuits capable of executing the above instructions (or instruction sets) individually or jointly. The electronic device 600 may also be a part of an integrated control system or a system manager, or may be configured as a portable electronic device that can be interconnected with a local or remote (e.g., via wireless transmission) interface.

[0113] In the electronic device 600, the processor 602 may include a central processing unit (CPU), a graphics processing unit (GPU), a programmable logic device, a dedicated processor system, a microcontroller, or a microprocessor. By way of example and not limitation, the processor may also include an analog processor, a digital processor, a microprocessor, a multi-core processor, a processor array, a network processor, and the like.

[0114] The processor 602 may run instructions or code stored in the memory 601, where the memory 601 may also store data. The instructions and data may also be sent and received via a network interface device over a network, where the network interface device may employ any known transmission protocol.

[0115] The memory 601 may be integrated with the processor 602. For example, RAM or flash memory may be disposed within an integrated circuit microprocessor or the like. In addition, the memory 601 may include a separate device, such as an external disk drive, a storage array, or other storage devices usable by any database system. The memory 601 and the processor 602 may be operatively coupled or may communicate with each other, for example, via an I / O port, a network connection, etc., such that the processor 602 can read files stored in the memory.

[0116] In addition, the electronic device 600 may further include a video display (such as a liquid crystal display) and a user interaction interface (such as a keyboard, a mouse, a touch input device, etc.). All components of the electronic device 600 may be connected to each other via a bus and / or a network.

[0117] According to an exemplary embodiment of the present disclosure, there is also provided a computer-readable storage medium including instructions, such as the memory 601 including instructions, and the above instructions may be executed by the processor 602 of the device 600 to complete the above method. Optionally, the computer-readable storage medium may be a ROM, a random access memory (RAM), a CD-ROM, a magnetic tape, a floppy disk, an optical data storage device, and the like.

[0118] According to an exemplary embodiment of the present disclosure, there may also be provided a computer program product, which includes a computer program / instructions, and when the computer program / instructions are executed by a processor, a method for image processing according to an exemplary embodiment of the present disclosure is implemented.

[0119] The above has been described with reference to Figures 1 to 6 an image processing method and apparatus according to an exemplary embodiment of the present disclosure. However, it should be understood that: Figure 5 the image processing apparatus and its units shown in Figure 6The electronic device shown is not limited to including the components shown above, but some components can be added or deleted as needed, and the above components can also be combined.

[0120] According to the image processing method and device of the present disclosure, by first obtaining an input image and a pixel classification map, a first transparency map, and a mask image corresponding to the input image, locally enhancing the input image based on the mask image, the pixel classification map, and the first transparency map to generate a second transparency map, and fusing the first transparency map and the second transparency map based on the mask image to generate a third transparency map of the input image, and then segmenting the input image based on the third transparency map, the speed and effect of image processing are improved.

[0121] Those skilled in the art will readily conceive of other embodiments of the present disclosure after considering the specification and practicing the invention disclosed herein. This application is intended to cover any variations, uses, or adaptations of the present disclosure that follow the general principles of the present disclosure and include known common knowledge or conventional technical means in the technical field not disclosed by the present disclosure. The specification and examples are only to be considered as exemplary, and the true scope and spirit of the present disclosure are pointed out by the following claims.

[0122] It should be understood that the present disclosure is not limited to the exact structures described above and shown in the drawings, and various modifications and changes can be made without departing from its scope. The scope of the present disclosure is only limited by the appended claims.

Claims

1. An image processing method, characterized in that, comprising: obtaining an input image, a pixel classification map, a first transparency map, and a mask image corresponding to the input image, wherein the pixel classification map indicates the category of each pixel in the input image, the category indicating whether each pixel is a transparent pixel or a non-transparent pixel, the mask image indicates the region for image processing, and the first transparency map is obtained from the downsampled image corresponding to the input image; performing local feature enhancement on the input image based on the mask image, the pixel classification map, and the first transparency map to generate a second transparency map; performing fusion processing on the first transparency map and the second transparency map based on the mask image to generate a third transparency map of the input image, wherein the resolution of the third transparency map is equal to the resolution of the input image; segmenting the input image based on the third transparency map.

2. The image processing method according to claim 1, characterized in that, the step of performing local feature enhancement on the input image based on the mask image, the pixel classification map, and the first transparency map comprises: respectively determining the regions to be optimized in the input image, the pixel classification map, and the first transparency map based on the mask image to obtain the region to be optimized in the input image, the region to be optimized in the pixel classification map, and the region to be optimized in the first transparency map; fusing the region to be optimized in the input image, the region to be optimized in the pixel classification map, and the region to be optimized in the first transparency map to obtain global information; performing local feature enhancement on the input image by inputting the global information into a preset local feature enhancement network.

3. The image processing method according to claim 1, characterized in that, the step of performing fusion processing on the first transparency map and the second transparency map based on the mask image comprises: determining the mask value of each pixel in the mask image as the weight of each pixel in the second transparency map; determining the difference between a preset value and the mask value of each pixel in the mask image as the weight of each pixel in the first transparency map; performing weighted summation on the first transparency map and the second transparency map based on the weight of each pixel in the first transparency map and the weight of each pixel in the second transparency map.

4. The image processing method according to claim 1, characterized in that, the step of obtaining the pixel classification map, the first transparency map, and the mask image corresponding to the input image comprises: obtaining the pixel classification map of the input image and the first transparency map of the input image; generating the mask image based on the first transparency map and the pixel classification map.

5. The image processing method according to claim 4, characterized in that, the step of obtaining the pixel classification map of the input image and the first transparency map of the input image comprises: obtaining a downsampled image corresponding to the input image; Obtain a pixel classification map of the input image and a first transparency map of the input image from the downsampled image, where the resolution of the first transparency map is less than the resolution of the input image.

6. The image processing method according to claim 5, wherein, the step of obtaining the first transparency map of the input image from the downsampled image includes: extracting semantic features of at least one level of the downsampled image; predicting the position of the object to be segmented and the boundary of the object to be segmented in the input image based on the semantic features of the at least one level; fusing the position of the object to be segmented and the boundary of the object to be segmented, and performing upsampling on the fusion result to obtain the first transparency map of the input image.

7. The image processing method according to claim 6, wherein, the step of obtaining the pixel classification map of the input image from the downsampled image includes: predicting the category of each pixel in the input image based on the semantic features of the at least one level; determining the pixel classification map of the downsampled image based on the category of each pixel in the input image; performing upsampling on the pixel classification map of the downsampled image to obtain the pixel classification map of the input image.

8. The image processing method according to claim 5, wherein, the step of obtaining the pixel classification map of the input image and the first transparency map of the input image from the downsampled image includes: obtaining the pixel classification map of the input image and the first transparency map of the input image from the downsampled image based on a preset global network.

9. The image processing method according to claim 8, wherein, the global network includes an encoder and a decoder, the encoder is configured to extract semantic features of at least one level of the downsampled image, the decoder is configured to perform feature reconstruction step by step using six levels, wherein the first three levels in the decoder are used to predict the position of the object to be segmented, and the last three levels in the decoder are used to predict the boundary of the object to be segmented.

10. The image processing method according to claim 2, wherein, the local feature enhancement network includes a local enhancement layer, and the local enhancement layer is used to enhance local features based on global information.

11. The image processing method according to claim 10, wherein, the local enhancement layer enhances local features by learning channel attention weights, and the channel attention weights are used to multiply with local features to encode foreground information and background information in global information into local features.

12. An image processing apparatus, wherein, comprising: an acquisition unit configured to acquire an input image and a pixel classification map, a first transparency map, and a mask image corresponding to the input image, where the pixel classification map indicates the category of each pixel in the input image, the category indicates whether each pixel is a transparent pixel or a non-transparent pixel, the mask image indicates the area for performing image processing, and the first transparency map is obtained from the downsampled image corresponding to the input image; A second transparency map generation unit, configured to perform local feature enhancement on the input image based on the mask image, the pixel classification map, and the first transparency map to generate a second transparency map; A third transparency map generation unit, configured to perform a fusion process on the first transparency map and the second transparency map based on the mask image to generate a third transparency map of the input image, where the resolution of the third transparency map is equal to the resolution of the input image; and An image segmentation unit, configured to segment the input image based on the third transparency map.

13. The image processing apparatus according to claim 12, wherein, The second transparency map generation unit is configured to: Based on the mask image, respectively determine the regions to be optimized in the input image, the pixel classification map, and the first transparency map, to obtain the region to be optimized in the input image, the region to be optimized in the pixel classification map, and the region to be optimized in the first transparency map; Fuse the region to be optimized in the input image, the region to be optimized in the pixel classification map, and the region to be optimized in the first transparency map to obtain global information; Perform local feature enhancement on the input image by inputting the global information into a preset local feature enhancement network.

14. The image processing apparatus according to claim 12, wherein, The third transparency map generation unit is configured to: Determine the mask value of each pixel in the mask image as the weight of each pixel in the second transparency map; Determine the difference between a preset value and the mask value of each pixel in the mask image as the weight of each pixel in the first transparency map; Perform a weighted sum of the first transparency map and the second transparency map based on the weight of each pixel in the first transparency map and the weight of each pixel in the second transparency map.

15. The image processing apparatus according to claim 12, wherein, The acquisition unit is configured to: Acquire the pixel classification map of the input image and the first transparency map of the input image; Generate the mask image based on the first transparency map and the pixel classification map.

16. The image processing apparatus according to claim 15, wherein, The acquisition unit is configured to: Acquire a downsampled image corresponding to the input image; Acquire the pixel classification map of the input image and the first transparency map of the input image from the downsampled image, where the resolution of the first transparency map is less than the resolution of the input image.

17. The image processing apparatus according to claim 16, wherein, The acquisition unit is configured to: Extract at least one level of semantic features of the downsampled image; Predict the position of the object to be segmented and the boundary of the object to be segmented in the input image based on the at least one level of semantic features; Fuse the position of the object to be segmented and the boundary of the object to be segmented, and perform upsampling on the fusion result to obtain the first transparency map of the input image.

18. The image processing apparatus according to claim 17, wherein, The acquisition unit is configured to: predict the class of each pixel in the input image based on the semantic features of the at least one level; determine a pixel classification map of the downsampled image based on the class of each pixel in the input image; upsample the pixel classification map of the downsampled image to obtain a pixel classification map of the input image.

19. The image processing apparatus according to claim 16, wherein, the acquisition unit is configured to: obtain a pixel classification map of the input image and a first transparency map of the input image from the downsampled image based on a preset global network.

20. The image processing apparatus according to claim 19, wherein, the global network includes an encoder and a decoder, the encoder is configured to extract semantic features of at least one level of the downsampled image, the decoder is configured to perform feature reconstruction step by step using six levels, wherein the first three levels in the decoder are used to predict the position of the object to be segmented, and the last three levels in the decoder are used to predict the boundary of the object to be segmented.

21. The image processing apparatus according to claim 13, wherein, the local feature enhancement network includes a local enhancement layer, and the local enhancement layer is used to enhance local features based on global information.

22. The image processing apparatus according to claim 21, wherein, the local enhancement layer enhances local features by learning channel attention weights, and the channel attention weights are used to multiply with local features to encode foreground information and background information in global information into local features.

23. An electronic device, wherein, comprises: a processor; a memory for storing instructions executable by the processor; wherein, the processor is configured to execute the instructions to implement the image processing method according to any one of claims 1 to 11.

24. A computer-readable storage medium storing a computer program, wherein, when the computer program is executed by a processor of an electronic device, the electronic device is caused to execute the image processing method according to any one of claims 1 to 11.

25. A computer program product comprising a computer program, wherein, when the computer program is executed by a processor, the image processing method according to any one of claims 1 to 11 is implemented.

Citation Information

Patent Citations

  • Interactive type image segmentation and fusion method based on grabucut algorithm

    CN107730528A

  • Natural image matting method based on deep learning

    CN111161277A