A method and device for portrait segmentation
The method employs a neural network with specific modules to segment human subjects from backgrounds on mobile devices, addressing the lack of real-time automatic background virtualization by reducing computational demands while maintaining accuracy.
Patent Information
- Application Number
- CN202010919568.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2020-09-04
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2040-09-04
AI Technical Summary
In the prior art, mobile devices cannot realize fully automatic real-time portrait segmentation and background blurring in application scenarios such as video chat, and users need to manually mark portraits, which cannot meet user needs.
The portrait segmentation network is adopted, including encoding module, pyramid pooling module and decoding module. The convolution layer and hollow convolution layer can be separated in depth reduce the calculation amount, realize portrait segmentation, and reduce the calculation needs.
While ensuring the accuracy of portrait segmentation, it reduces the amount of computing, reduces the computing power requirements of terminal devices, and realizes fully automatic real-time portrait segmentation and background blurring.
Smart Images

Figure CN114219941B_ABST
Abstract
Description
Technical Field
[0001] This application belongs to the field of digital image processing, and particularly relates to a method and device for portrait segmentation. Background Art
[0002] With the development of technology, more and more people rely on mobile devices such as mobile phones to take photos and videos. However, due to factors such as the hardware performance of mobile devices, it is unable to meet the user's requirements for real-time background blurring and replacement of the captured content in application scenarios such as video chats and video conferences.
[0003] In the prior art, when using a mobile device to achieve background blurring for an image or video containing a portrait, the user needs to manually mark the portrait in order to achieve background blurring, but in this way, it is impossible to achieve fully automatic processing of the captured content and meet the user's needs in application scenarios such as video chats. Summary of the Invention
[0004] Embodiments of this application provide a method and device for portrait segmentation, which can segment the portrait in an image or video based on a portrait segmentation network, and solve the problem in the prior art that it is impossible to achieve fully automatic real-time portrait segmentation in a mobile device with limited computing power to blur and replace the background.
[0005] In a first aspect, embodiments of this application provide a method for portrait segmentation, including:
[0006] Obtain a raw portrait image, where the picture of the raw portrait image contains a portrait;
[0007] Input the raw portrait image into a portrait segmentation network for processing, and output a portrait mask image; wherein, the portrait segmentation network includes an encoding module, a pyramid pooling module, and a decoding module;
[0008] Divide the raw portrait image into a portrait image and a background image according to the portrait mask image.
[0009] In a possible implementation manner of the first aspect, after dividing the raw portrait image into a portrait image and a background image according to the portrait mask image, the portrait segmentation method further includes:
[0010] Perform special processing on the background image according to the user operation to generate a new background image;
[0011] Combine the portrait image and the new background image to generate a new portrait image, and display the new portrait image.
[0012] Exemplarily, the above special processing can specifically be blurring, replacement, adding filters, or other image processing methods. The specific method still needs to be determined according to the user's needs (learned by receiving user operations). The image generated after the background image is specially processed is the new background image.
[0013] It should be understood that the above acquisition of the original portrait image of the user includes: acquiring the original portrait video of the user; identifying each frame image in the original portrait video as the original portrait image. That is, the method provided in this application can not only perform portrait segmentation on an image, but also perform portrait segmentation on a video composed of multiple frame images to achieve real-time portrait segmentation. Subsequently, after dividing the original portrait image into a portrait image and a background image according to the portrait mask image, the portrait segmentation method further includes: generating a portrait video according to the portrait images corresponding to all the frame images; generating a background video according to the background images corresponding to all the frame images.
[0014] In a second aspect, an embodiment of the present application provides a portrait segmentation device, including:
[0015] A module for acquiring the original portrait image, which is used to acquire the original portrait image, and the picture of the original portrait image contains a portrait;
[0016] A portrait segmentation network module, which is used to input the original portrait image into a portrait segmentation network for processing and output a portrait mask image; wherein, the portrait segmentation network module includes an encoding module, a pyramid pooling module, and a decoding module;
[0017] A portrait segmentation module, which is used to divide the original portrait image into a portrait image and a background image according to the portrait mask image.
[0018] In a third aspect, an embodiment of the present application provides a terminal device, which includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the computer program is executed by the processor, it implements the method described in any item of the first aspect above.
[0019] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, it implements the method described in any item of the first aspect above.
[0020] In a fifth aspect, an embodiment of the present application provides a computer program product. When the computer program product runs on a terminal device, it causes the terminal device to execute the method described in any item of the first aspect above.
[0021] It can be understood that the beneficial effects of the above second aspect to the fifth aspect can refer to the relevant descriptions in the first aspect above, and will not be elaborated here.
[0022] The beneficial effects of the embodiments of the present application compared with the prior art are as follows:
[0023] The method provided by the present application, compared with the prior art, when constructing a portrait segmentation network, sets an encoding module, a pyramid pooling module, and a decoding module, reducing the computational amount when the portrait segmentation network obtains a portrait mask image; when segmenting a portrait in an image or video based on the portrait segmentation network, while ensuring the accuracy of portrait segmentation, the computational amount is reduced, thereby reducing the computational power requirements of the terminal device equipped with the portrait segmentation network. BRIEF DESCRIPTION OF THE DRAWINGS
[0024] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0025] Figure 1 is the implementation flowchart of the method provided by the first embodiment of the present application;
[0026] Figure 2 is the implementation flowchart of the method provided by the second embodiment of the present application;
[0027] Figure 3 is the structural schematic diagram of the portrait segmentation network provided by an embodiment of the present application;
[0028] Figure 4 is the implementation flowchart of the method provided by the third embodiment of the present application;
[0029] Figure 5 is the structural schematic diagram of the pyramid pooling module provided by an embodiment of the present application;
[0030] Figure 6 is the implementation flowchart of the method provided by the fourth embodiment of the present application;
[0031] Figure 7 is the structural schematic diagram of the decoding unit provided by an embodiment of the present application;
[0032] Figure 8 is the implementation flowchart of the method provided by the fifth embodiment of the present application;
[0033] Figure 9 is the structural schematic diagram of the conversion unit provided by an embodiment of the present application;
[0034] Figure 10 is the implementation flowchart of the method provided by the sixth embodiment of the present application;
[0035] Figure 11 is the implementation flowchart of the method provided in the seventh embodiment of this application;
[0036] Figure 12 is the visualization diagram of the masked ground truth image and the boundary ground truth image provided in an embodiment of this application;
[0037] Figure 13 is the schematic diagram of the application scenario provided in an embodiment of this application;
[0038] Figure 14 is the schematic structural diagram of the device provided in an embodiment of this application;
[0039] Figure 15 is the schematic structural diagram of the terminal device provided in an embodiment of this application. Detailed implementation manners
[0040] In the following description, specific details such as specific system structures and technologies are presented for the purpose of illustration rather than limitation, so as to thoroughly understand the embodiments of this application. However, those skilled in the art should understand that this application can also be implemented in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, devices, circuits, and methods are omitted to avoid unnecessary details from interfering with the description of this application.
[0041] It should be understood that when used in the specification and the appended claims of this application, the term "comprising" indicates the presence of the described features, wholes, steps, operations, elements, and / or components, but does not exclude the presence or addition of one or more other features, wholes, steps, operations, elements, components, and / or their combinations.
[0042] It should also be understood that the term "and / or" as used in the specification and the appended claims of this application refers to any combination and all possible combinations of one or more of the associated listed items, and includes these combinations.
[0043] As used in the specification and the appended claims of this application, the term "if" can be interpreted as "when", "once", "in response to determining", or "in response to detecting" according to the context. Similarly, the phrase "if determined" or "if [the described condition or event] is detected" can be interpreted as meaning "once determined", "in response to determining", "once [the described condition or event] is detected", or "in response to detecting [the described condition or event]" according to the context.
[0044] In addition, in the description of the specification and the appended claims of this application, the terms "first", "second", "third", etc. are only used for distinguishing descriptions and cannot be understood as indicating or implying relative importance.
[0045] References to "one embodiment" or "some embodiments" or the like described in the specification of this application mean that a particular feature, structure, or characteristic described in connection with that embodiment is included in one or more embodiments of this application. Thus, statements such as "in one embodiment", "in some embodiments", "in other some embodiments", "in still other embodiments", etc. that appear in different places in this specification do not necessarily all refer to the same embodiment, but rather mean "one or more but not all embodiments", unless otherwise specifically emphasized. The terms "comprising", "including", "having" and their variants all mean "including but not limited to", unless otherwise specifically emphasized.
[0046] Regarding the term explanations of this application, a brief description is given here. "Feature information" includes the correlation relationships among the individual pixels within a feature image. "Size" refers to the pixels (height and width) of the feature image. "Channel" refers to the number of layers of the feature image. "Feature fusion" means: performing feature fusion on feature image A and feature image B to extract the feature information contained in both. Specifically: if the sizes of feature image A and feature image B are the same, then the two are directly concatenated in channels to obtain a feature fusion image with the same size as the two and a channel number equal to the sum of the two; if the sizes of feature image A and feature image B are different, then one of the feature images is processed to make its size the same as that of the other feature image. Specifically, it can be a bilinear interpolation operation on the feature image with the smaller size to obtain a feature image with the same size as the feature image with the larger size. "Channel adjustment convolution processing" means: performing point convolution processing with a convolution kernel of 1*1 on a target feature image, mainly used to adjust the channel number of the target feature image. Specifically, by adjusting the convolution parameters in the point convolution processing with a convolution kernel of 1*1 (the convolution parameter is specifically depth_multiplier, which is used to control the number of convolution output channels in the depth direction of each input channel, and the value of this convolution parameter is the ratio of the total output channel number to the total input channel number), the channel number of the target feature image can be adjusted, including increasing, decreasing, or keeping the channel number unchanged. "Depthwise separable convolution processing" means performing depth convolution processing on a target feature image and then performing point convolution processing on the result of the depth convolution processing. Specifically, by adjusting the convolution parameters of the depthwise separable convolution processing, the size and channel number of the target feature image can be adjusted.
[0047] In the embodiments of this application, the execution subject of the process is a terminal device. The terminal device includes, but is not limited to: devices such as servers, computers, smartphones, and tablet computers that can execute the portrait segmentation method provided in this application. Preferably, the terminal device is a mobile terminal, and the mobile terminal can acquire the original portrait image of the user. Figure 1The implementation flowchart of the method provided by the first embodiment of the present application is shown and described in detail as follows:
[0048] In S101, an original portrait image is obtained, and the portrait in the original portrait image is included in the picture.
[0049] In this embodiment, the portrait in the original portrait image is included in the picture. The above-mentioned obtaining of the original portrait image can specifically be: obtaining the original portrait image through the camera of the terminal device.
[0050] It should be understood that the obtaining of the original portrait image includes: obtaining an original portrait video; identifying each frame image in the original portrait video as the above-mentioned original portrait image. That is, the method provided by this embodiment can not only perform portrait segmentation on images, but also perform portrait segmentation on each frame image in the video to achieve real-time portrait segmentation of the video.
[0051] In S102, the original portrait image is input into a portrait segmentation network for processing, and a portrait mask image is output.
[0052] In this embodiment, the portrait segmentation network includes an encoding module, a pyramid pooling module, and a decoding module. The encoding module can be a preset convolutional neural network. Compared with a general convolutional neural network, the encoding module uses a depthwise separable convolutional layer instead of a standard convolutional layer to reduce the amount of calculation when extracting the feature information of the original portrait image; the pyramid pooling module is used to obtain global information about the original portrait image according to the feature information to improve the accuracy when outputting the subsequent portrait mask image; the decoding module is used to generate the portrait mask image according to the feature information and the global information.
[0053] In a possible implementation manner, the above-mentioned inputting the original portrait image into the portrait segmentation network for processing and outputting the portrait mask image can specifically be: inputting the original portrait image into the encoding module for processing to output a feature information image; inputting the feature information image into the pyramid pooling module for processing to output a global information image; inputting the feature information and the global information image into the decoding module for processing to output the portrait mask image.
[0054] In S103, the original portrait image is divided into a portrait image and a background image according to the portrait mask image.
[0055] In this embodiment, the portrait mask image contains the classification information of each pixel in the original portrait image. The original portrait image is divided into a portrait image and a background image according to the portrait mask image. Specifically, each pixel in the original portrait image is divided into a portrait pixel or a background pixel according to the portrait mask image; a portrait image is generated based on all the portrait pixels; and a background image is generated based on all the background pixels. It should be understood that dividing the original portrait image into a portrait image and a background image realizes portrait segmentation.
[0056] In a possible implementation manner, the portrait mask image contains the probability of each pixel in the original portrait image being a portrait pixel. The above-mentioned step of dividing each pixel in the original portrait image into a portrait pixel or a background pixel according to the portrait mask image can be specifically as follows: if the probability of the pixel being a portrait pixel is greater than or equal to a preset threshold (exemplarily, the preset threshold is 0.5), then the pixel is identified as a portrait pixel; if the probability of the pixel being a portrait pixel is less than the preset threshold (exemplarily, the preset threshold is 0.5), then the pixel is identified as a background pixel.
[0057] In this embodiment, when constructing the portrait segmentation network, by setting an encoding module, a pyramid pooling module, and a decoding module, the structure of the portrait segmentation network is optimized compared with the neural network in the prior art. A depthwise separable convolutional layer is used instead of a standard convolutional layer in the encoding module. Since the convolutional kernel of the depthwise separable convolutional layer is often smaller than that of the standard convolutional layer, the computational amount of the encoding module is reduced, and then the computational amount of the portrait segmentation network is reduced; when the portrait in the image or video is segmented based on the portrait segmentation network, while reducing the computational amount, the accuracy of portrait segmentation is ensured, thereby reducing the computational power requirement of the device carrying the portrait segmentation network.
[0058] Figure 2 The flowchart of the method provided in the second embodiment of the present application is shown. Refer to Figure 2 , relative to the above Figure 1 described embodiment, the method S102 provided in this embodiment includes S201 to S204, which are specifically described in detail as follows:
[0059] Further, the above-mentioned step of inputting the original portrait image into the portrait segmentation network for processing and outputting the portrait mask image includes:
[0060] In S201, the original portrait image is input into the encoding module, and the original portrait image is subjected to multiple convolutional processes using the depthwise separable convolutional algorithm corresponding to the encoding module, and N downsampled feature images are output.
[0061] In this embodiment, N is an integer greater than 2. Generally, the value range of N is from 4 to 6, and the preferred value of N is 5. This value is used to balance the computational amount and the recognition accuracy of the human portrait segmentation network. Compared with a general convolutional neural network (such as MobileNetV2), a depthwise separable convolutional layer is adopted in the encoding module to replace the standard convolutional layer. Since the convolutional kernel of the depthwise separable convolutional layer is usually smaller than that of the standard convolutional layer, the computational amount for extracting feature information from the original human portrait image is reduced compared with a general convolutional neural network. The depthwise separable convolutional layer includes at least one depthwise convolutional layer (depthwise conv) and at least one pointwise convolutional layer (pointwise conv).
[0062] In this embodiment, the original human portrait image is input into the encoding module, and the depthwise separable convolutional algorithm corresponding to the encoding module is used to perform multiple convolutional processes on the original human portrait image, and N downsampled feature images are output. Specifically, it can be: perform a channel adjustment convolutional process on the original human portrait image to obtain a channel-expanded image; input the channel-expanded image into a series network composed of multiple encoding units of the encoding module to obtain multiple encoded result images. The encoding unit includes at least one depthwise separable convolutional layer. Optionally, the encoding unit may further include a downsampling layer, and the ratio of the input image size to the output image size of the downsampling layer conforms to a preset downsampling ratio; generally, the input image size is larger than the output image size, and the downsampling ratio between the two can be 2; select N of the encoded result images as the N downsampled feature images. Perform a channel adjustment convolutional process on the original human portrait image to obtain a channel-expanded image. Specifically, by presetting the convolutional parameters of the channel adjustment convolutional process, the number of channels of the original human portrait image can be increased to obtain the channel-expanded image with an increased number of channels.
[0063] Figure 3 FIG. shows a schematic structural diagram of a human portrait segmentation network provided by an embodiment of the present application. In a possible implementation manner, as Figure 3 shown, the encoding module of the human portrait segmentation network includes 5 encoding units, that is, the above-mentioned N = 5. Among them, encoding units 1, 2, 3, and 4 all include downsampling layers, and encoding unit 5 does not include a downsampling layer. Each encoding unit in the figure outputs a corresponding downsampled feature image. Input the original human portrait image into the encoding module for processing, so that the original human portrait image sequentially passes through the channel adjustment convolution, encoding unit 1, encoding unit 2, encoding unit 3, encoding unit 4, and encoding unit 5 in the figure, and output the first downsampled feature image, the second downsampled feature image, the third downsampled feature image, the fourth downsampled feature image, and the fifth downsampled feature image.
[0064] It should be understood that the coding unit 5 shown in the figure, that is, the fifth coding unit, can be referred to here when explaining all the units followed by numbers in the figures in this application; the above-mentioned first downsampled feature image, that is, the first downsampled feature image, can be referred to here when explaining all the image names prefixed with numbers in this application. For example, the Nth downsampled feature image or the Nth first downsampled feature image refers to the same thing.
[0065] In S202, the Nth downsampled feature image is input into the pyramid pooling module for processing, and the Nth upsampled feature image is output.
[0066] In this embodiment, compared with the pooling layer of the existing neural network, a dilated convolutional layer is used instead of the pooling layer in the pyramid pooling module. Specifically, the above-mentioned input of the Nth downsampled feature image into the pyramid pooling module for processing and outputting the Nth upsampled feature image can be: performing dilated convolution (dilatedconv) processing on the Nth downsampled feature image to obtain a dilated convolutional feature image; adjusting the size of the dilated convolutional image to make the size of the dilated convolutional feature image the same as that of the Nth downsampled feature image; and performing channel splicing on the dilated convolutional feature image and the Nth downsampled feature image to obtain the Nth upsampled feature image.
[0067] In a possible implementation manner, as Figure 3 shown, N is 5, and the fifth downsampled feature image is input into the pyramid pooling module for processing, and the fifth upsampled feature image is output.
[0068] In S203, the Nth upsampled feature image and the N - 1 downsampled feature images other than the Nth downsampled feature image among the N downsampled feature images are input into the decoding module for processing, and the first upsampled feature image is output.
[0069] In this embodiment, the decoding module includes N - 1 decoding units. The (M - 1)th decoding unit is used to determine the (M - 1)th upsampled feature image according to the Mth upsampled feature image and the (M - 1)th downsampled feature image, where M is an integer less than or equal to N and greater than 0. The size of the (M - 1)th downsampled feature image is the same as that of the (M - 1)th upsampled feature image.
[0070] In this embodiment, inputting the Nth upsampled feature image and the N - 1 downsampled feature images (excluding the Nth downsampled feature image) among the N downsampled feature images into the decoding module for processing to output the first upsampled feature image may specifically include: At the beginning, when the value of M is N, input the Nth upsampled feature image and the (N - 1)th downsampled feature image into the (N - 1)th decoding unit for processing to obtain the (N - 1)th upsampled feature image; and so on. Finally, when the value of M is 2, input the second upsampled feature image and the first downsampled feature image into the first decoding unit for processing to obtain the first upsampled feature image.
[0071] In a possible implementation manner, as Figure 3 shown, N is 5. Input the fifth upsampled feature image and the fourth downsampled feature image into the illustrated decoding unit 4 (i.e., the fourth decoding unit) for processing to output the fourth upsampled feature image... and so on. Finally, the illustrated decoding unit 1 (i.e., the first decoding unit) outputs the first upsampled feature image.
[0072] In S204, perform image masking processing on the first upsampled feature image to obtain a portrait mask image.
[0073] In this embodiment, the above-mentioned performing image masking processing on the first upsampled feature image to obtain a portrait mask image may specifically be: Refer to Figure 3 , perform a single-channel adjustment convolution processing on the first upsampled feature image. By presetting the convolution parameters of the channel adjustment convolution processing, the number of channels of the first upsampled feature image can be reduced to obtain a channel-reduced image with a reduced number of channels; perform a bilinear interpolation operation on the channel-reduced image to make the size of the channel-reduced image consistent with the above-mentioned original portrait image, that is, to enlarge the size of the channel-reduced image to obtain the portrait mask image.
[0074] In this embodiment, the encoding module uses depthwise separable convolutional layers instead of standard convolutional layers, which helps to reduce the computational amount of the portrait segmentation network, improve the output efficiency, and reduce the computational power requirements of the terminal device (preferably a mobile device) implementing the embodiments of the present application; the pyramid pooling module uses dilated convolutional layers instead of pooling layers, which can enable the portrait segmentation network to obtain rich global information of the original portrait image to improve the accuracy of subsequent portrait segmentation; the decoding module includes a plurality of decoding units corresponding to the encoding module to implement obtaining the portrait mask image from the original portrait image, which is convenient for subsequent portrait segmentation.
[0075] Figure 4 shows the implementation flowchart of the method provided in the third embodiment of the present application. Refer to Figure 4, relative to the above Figure 2 In the embodiment provided, method S202 includes S2021 to S2023, which are specifically described in detail as follows:
[0076] Further, inputting the Nth downsampled feature image into the pyramid pooling module for processing and outputting the Nth upsampled feature image includes:
[0077] In S2021, perform atrous convolution processing on the Nth downsampled feature image respectively using a plurality of preset convolution kernels with different scales to obtain a plurality of atrous convolution feature images.
[0078] In this embodiment, the above-mentioned plurality of atrous convolution feature images correspond one by one to the above-mentioned plurality of preset convolution kernels with different scales; this atrous convolution processing is used to increase the receptive field of the Nth downsampled feature image to improve the ability of the portrait segmentation network to obtain more feature information in the Nth downsampled feature image. When performing this atrous convolution processing on the Nth downsampled feature image, using a plurality of convolution kernels with different scales is to obtain different atrous convolution feature images, so as to enrich the feature information of the subsequent obtained Nth upsampled feature image. The number of the above-mentioned plurality of preset convolution kernels with different scales is preferably 4 to balance the beneficial effects in terms of enriching the feature information of the Nth upsampled feature image and reducing the computational amount.
[0079] Figure 5 shows a schematic structural diagram of a pyramid pooling module provided in an embodiment of the present application. In a possible implementation manner, as Figure 5 shown, perform atrous convolution processing on the Nth downsampled feature image respectively using four convolution kernels of 1*1, 2*2, 3*3, and 6*6 to obtain 4 atrous convolution feature images.
[0080] In S2022, perform convolution processing on each atrous convolution feature image respectively to obtain a plurality of to-be-fused feature images corresponding one by one to the plurality of atrous convolution feature images.
[0081] In this embodiment, performing convolution processing on each atrous convolution feature image respectively to obtain a plurality of to-be-fused feature images corresponding one by one to the above-mentioned plurality of atrous convolution feature images specifically includes: performing a channel adjustment convolution processing on the atrous convolution feature image. By presetting the convolution parameters of the channel adjustment convolution processing, the channels of the atrous convolution feature image can be reduced to obtain the to-be-fused feature image with reduced channels. Preferably, the sum of the channels of all the to-be-fused feature images is the same as that of the Nth downsampled feature image, and the number of channels of each to-be-fused feature image is the same.
[0082] In a possible implementation manner, as Figure 5As shown, convolution processing is respectively performed on four dilated convolution feature images to obtain four feature images to be fused, and the number of channels of each feature image to be fused is one-fourth of that of the Nth downsampled feature image.
[0083] In S2023, multiple feature images to be fused are subjected to feature fusion to obtain a pooled feature image.
[0084] In this embodiment, the above-mentioned multiple feature images to be fused are subjected to feature fusion to obtain a pooled feature image, which specifically may be: respectively performing bilinear interpolation operations on the above-mentioned multiple feature images to be fused to obtain multiple spliced feature images to be stitched that are in one-to-one correspondence with the above-mentioned multiple feature images to be fused and have the same size, and the size of the spliced feature image to be stitched is the same as that of the Nth downsampled feature image; performing channel stitching on all the spliced feature images to be stitched to obtain a pooled feature image with the same number of channels as the Nth downsampled feature image.
[0085] In a possible implementation manner, as Figure 5 shown, performing feature fusion on four feature images to be fused to obtain a pooled feature image, which specifically may be: respectively performing bilinear interpolation operations on each feature image to be fused to obtain four spliced feature images to be stitched with the same size as the Nth downsampled feature image; performing channel stitching on the four spliced feature images to be stitched to obtain a pooled feature image with the same size and the same number of channels as the Nth downsampled feature image.
[0086] In S2024, the pooled feature image and the Nth downsampled feature image are subjected to channel stitching to obtain the Nth upsampled feature image.
[0087] In this embodiment, the number of channels of the Nth upsampled feature image is twice that of the Nth downsampled feature image.
[0088] In this embodiment, in the pyramid pooling module, convolution processing is performed on the Nth downsampled feature image using multiple convolution kernels with different scales, so as to extract richer feature information; subsequently, the pooled feature image and the Nth downsampled feature image are subjected to channel stitching, and the obtained Nth upsampled feature image contains richer global information and aggregates the context information of different-scale convolutional layers, so as to improve the accuracy of portrait segmentation subsequently.
[0089] Figure 6 shows the implementation flowchart of the method provided in the fourth embodiment of the present application. Refer to Figure 6 , relative to the above Figure 2 described embodiment, the method S203 provided in this embodiment includes S2031 to S2033, which are specifically described in detail as follows:
[0090] Further, determining the (M - 1)-th upsampled feature image based on the M-th upsampled feature image and the (M - 1)-th downsampled feature image includes:
[0091] In this embodiment, the decoding unit includes an upsampling layer, a feature fusion layer, and a conversion unit.
[0092] In S2031, the M-th upsampled feature image is input into the upsampling layer of the (M - 1)-th decoding unit, and an upsampled feature mapping is performed on the M-th upsampled feature image to output the (M - 1)-th initial feature image.
[0093] In this embodiment, the ratio of the input image size to the output image size of the upsampling layer conforms to a preset upsampling ratio; among them, the input image size is generally smaller than the output image size, and the downsampling ratio of the two can be 0.5; the upsampling ratio is the reciprocal of the above downsampling ratio; the above-mentioned step of inputting the M-th upsampled feature image into the upsampling layer of the (M - 1)-th decoding unit, performing an upsampled feature mapping on the M-th upsampled feature image, and outputting the (M - 1)-th initial feature image can specifically be: performing a bilinear interpolation operation on the M-th upsampled feature image to obtain the (M - 1)-th initial feature image with an enlarged size.
[0094] In a possible implementation manner, as Figure 3 shown, it should be noted that the fourth decoding unit does not include the above upsampling layer. There are 4 downsampling layers in the encoding module shown in the figure. The sizes of the original portrait image and the source portrait image should be the same. Therefore, the decoding module should have 3 upsampling layers to offset the size changes of the 3 downsampling layers (the bilinear interpolation in the figure offsets the size change of the remaining 1 downsampling layer); the encoding unit 5 shown in the figure does not include a downsampling layer, that is, the sizes of the fourth downsampled feature image and the fifth downsampled feature image are the same, that is, the sizes of the fourth downsampled feature image and the fifth upsampled feature image are the same. Therefore, the fourth decoding unit does not include the above upsampling layer. It should be understood that whether the fourth decoding unit in the figure includes an upsampling layer depends on the total number of downsampling layers in the encoding module.
[0095] At this time, if M is 5, the step of inputting the fifth upsampled feature image into the above upsampling layer of the fourth decoding unit, performing an upsampled feature mapping on the fifth upsampled feature image, and outputting the fourth initial feature image is not executed. Instead, the fifth upsampled feature image is identified as the fourth initial feature image for subsequent steps.
[0096] In S2032, the (M-1)th initial feature image and the (M-1)th downsampled feature image are input into the feature fusion layer of the (M-1)th decoding unit, and the (M-1)th initial feature image and the (M-1)th downsampled feature image are subjected to feature fusion to obtain the (M-1)th upsampled feature fusion image.
[0097] In this embodiment, the (M-1)th initial feature image has the same size as the (M-1)th downsampled feature image. Therefore, the above-mentioned feature fusion of the (M-1)th initial feature image and the (M-1)th downsampled feature image to obtain the (M-1)th upsampled feature fusion image can specifically be to perform channel splicing on the (M-1)th initial feature image and the (M-1)th downsampled feature image to obtain the (M-1)th upsampled feature fusion image with the number of channels being the sum of the two and the size being the same as the two.
[0098] Figure 7 The schematic diagram of the decoding unit structure provided by an embodiment of the present application is shown. Refer to Figure 3 , Figure 7 The shown schematic diagram of the decoding unit structure is only used to describe Figure 3 the decoding unit including the upsampling layer in Figure 3 and the decoding unit without the upsampling layer in Figure 7 can be referred to, and the difference is only that it does not include the upsampling layer.
[0099] In a possible implementation manner, as Figure 7 shown, the Mth upsampled feature image is input into the upsampling layer in the shown decoding unit for processing, and the (M-1)th downsampled feature image is input into the feature fusion layer in the shown decoding unit for processing. Finally, the decoding unit outputs the (M-1)th upsampled feature image.
[0100] In S2033, the (M-1)th upsampled feature fusion image is input into the conversion unit of the (M-1)th decoding unit for processing, and the (M-1)th upsampled feature image is output.
[0101] In this embodiment, the above-mentioned input of the (M-1)th upsampled feature fusion image into the conversion unit of the (M-1)th decoding unit for processing and outputting the (M-1)th upsampled feature image can specifically be: performing at least one depthwise separable convolution processing on the (M-1)th upsampled feature fusion image to obtain the (M-1)th upsampled feature image.
[0102] In this embodiment, particularly, when outputting the upsampled feature image, by performing feature fusion with the downsampled feature image, the feature information of the output upsampled feature image can be enriched, including the relevance with the downsampled feature information of each layer, so as to improve the accuracy of portrait segmentation.
[0103] Figure 8 Shows the implementation flowchart of the method provided in the fifth embodiment of the present application. Refer to Figure 8 , compared with the above Figure 7 described embodiment, the method S2033 provided in this embodiment includes S801 to S803, which are specifically described in detail as follows:
[0104] In this embodiment, the conversion unit includes at least one depthwise separable convolutional layer, at least one channel adjustment convolutional layer, and at least one feature fusion block. Figure 9 Shows the schematic structural diagram of the conversion unit provided in an embodiment of the present application. In a possible implementation manner, as Figure 9 shown in (a) of
[0105] further, processing the conversion unit that inputs the (M-1)th upsampled feature fusion image into the (M-1)th decoding unit, and outputting the (M-1)th upsampled feature image includes:
[0106] In S801, input the (M-1)th upsampled feature fusion image into at least one depthwise separable convolutional layer for processing, and output the (M-1)th image to be converted.
[0107] In a possible implementation manner, refer to Figure 9 shown in (a) of
[0108] In S802, input the (M-1)th upsampled feature fusion image into at least one channel adjustment convolutional layer for processing, and output the (M-1)th channel adjustment image.
[0109] In a possible implementation manner, refer to Figure 9 shown in (a) of
[0110] In S803, the (M-1)th image to be converted and the (M-1)th channel-adjusted image are input into at least one feature fusion block for feature fusion, and the (M-1)th upsampled feature image is output.
[0111] In a possible implementation, referring to Figure 9 (a) in, the (M-1)th image to be converted and the (M-1)th channel-adjusted image are input into the illustrated feature fusion block for feature fusion, and the (M-1)th upsampled feature image is output.
[0112] Specifically, the convolution kernel of each Figure 9 depth convolution layer shown in is 3*3, and the convolution kernel of each Figure 9 point convolution layer shown in is 1*1. Each Figure 9 convolution layer shown in includes a RELU activation function. It should be understood that after each Figure 9 convolution layer, there is a step of batch normalization (BatchNormalization, BN) to avoid the problem of internal covariate shift in the human portrait segmentation network. For specific details, reference can be made to the relevant descriptions of the prior art, which will not be elaborated here.
[0113] In another possible implementation, as shown in Figure 9 (b) in, the conversion unit includes four depthwise separable convolution layers, two channel-adjusting convolution layers, and two feature fusion blocks; the implementation of S2033 can be specifically referred to the steps of the above possible implementation. The difference is that here the implementation of S2033 is to repeat the steps of the above possible implementation twice, using the output of the first execution of the steps of the above possible implementation as the input of the second execution of the steps of the above possible implementation, and using the output of the second execution of the steps of the above possible implementation as the output of the conversion unit.
[0114] In this embodiment, by setting a depthwise separable convolution layer and a channel-adjusting convolution layer in the conversion unit, and performing feature fusion on the outputs of the two through a feature fusion block, a feature image with more internal connections (i.e., feature information) of each pixel point is obtained. Setting more depthwise separable convolution layers, channel-adjusting convolution layers, and feature fusion blocks can improve the robustness and accuracy of the human portrait segmentation model, but it will also increase the computational amount. Therefore, when setting the conversion unit, the requirements of human portrait segmentation should be considered. Through the two possible implementation methods detailed in this embodiment, the human portrait segmentation network can run normally on a mobile device.
[0115] Figure 10 shows the implementation flowchart of the method provided in the sixth embodiment of the present application. Refer to Figure 10, relative to any of the above embodiments, the method provided in this embodiment includes S10-a to S10-e, which are specifically described in detail as follows:
[0116] Further, before inputting the original portrait image into the portrait segmentation network for processing and outputting the portrait mask image, the portrait segmentation method further includes:
[0117] In this embodiment, generally, the steps of S10-a to S10-e in this embodiment can be executed on a training device outside the terminal device. The portrait segmentation network is trained on the training device through the steps of S10-a to S10-e, and then the trained portrait segmentation network is sent to the terminal device so that the terminal device can carry the portrait segmentation network and operate normally. The portrait segmentation network in this embodiment has a relatively low computing power requirement for the carrying device.
[0118] In S10-a, a training image set is obtained.
[0119] In this embodiment, the training image set includes a number of training images, and the picture of the training image contains a portrait. Generally, the training image contains the portrait of the target object and the background outside the portrait.
[0120] In a possible implementation manner, in order to improve the generalization of the portrait segmentation network, data augmentation is performed when obtaining the training image set. Specifically, the original training image is obtained; the original training image is subjected to deformation augmentation processing to obtain a deformed training image; the original training image is subjected to texture augmentation processing to obtain a textured training image; the original training image, the deformed training image, and the textured training image are all recognized as the training images in the training image set. The deformation augmentation processing includes: randomly flipping the image up, down, forward, and backward, randomly scaling the image, randomly trimming the image, and randomly rotating the image, etc.; it should be understood that the deformed training image after the deformation augmentation processing has the same size as the original training image. After operations that affect the image size such as randomly scaling the image or randomly trimming the image, the image size needs to be adjusted to make the deformed training image have the same size as the original training image. The texture augmentation processing includes: randomly adding image weather effects (such as rain, fog, sunny, cloudy), randomly adjusting the image color, and randomly adjusting the image contrast, etc.; it should be understood that the texture augmentation processing should not affect the portrait area part of the original training image, that is, the above texture augmentation processing is only applied to the background area part of the original training image, and the texture augmentation processing is realized by changing the pixel values in the background area part of the original training image.
[0121] In S10-b, the training image is configured with an associated mask ground truth image.
[0122] In this embodiment, in order to train the human portrait segmentation network, it is necessary to perform data annotation on each training image in the training image set, that is, to mark the human portrait area and the background area on each training image.
[0123] In a possible implementation manner, preferably, the above training image is configured with an associated mask ground-truth image. The specific implementation includes: in the training image, the pixel values within the marked human portrait area are set to 255, and the pixel values within the marked background area are set to 0, to obtain the mask ground-truth image, which is associated with the training image.
[0124] In another possible implementation manner, optionally, a mask ground-truth image is configured for each training image in the training image set to obtain a plurality of mask ground-truth images corresponding one by one to the above-mentioned plurality of training images. Specifically, taking one training image as an example for illustration, in the training image, the pixel values within the marked human portrait area are set to 0, and the pixel values within the marked background area are set to 255, to obtain the mask ground-truth image, which is associated with the training image.
[0125] In S10-c, a number of training images are input into the human portrait segmentation network for processing, and prediction mask images corresponding one by one to the number of training images are output.
[0126] In this embodiment, since the implementation manner of S10-c is completely the same as that of S102 in the above Figure 1 described embodiment, the specific description can refer to the relevant description of S102 and will not be elaborated here.
[0127] In S10-d, the network loss value of the human portrait segmentation network is calculated according to the prediction mask image and the mask ground-truth image.
[0128] In this embodiment, the prediction mask image is any one of the above-mentioned number of prediction mask images, and the mask ground-truth image is any one of the above-mentioned number of mask ground-truth images.
[0129] In this embodiment, taking one of the training images as an example for illustration, the training image corresponds to the predicted mask image and the ground truth mask image. Using the training image as the input, the predicted mask image as the predicted value (the output of the human segmentation network), and the ground truth mask image as the ground truth label (the ideal output of the human segmentation network), the human segmentation network is trained so that when the human segmentation network takes the training image as the input, the output is as close as possible to the ground truth mask image. Specifically, according to the difference between the predicted mask image and the ground truth mask image, the network loss value of the human segmentation network is calculated, so that subsequently, the weight parameters of the human segmentation network can be adjusted according to the network loss value, so that after adjustment, when the human segmentation network takes the training image as the input, the output is as close as possible to the ground truth mask image.
[0130] In a possible implementation manner, the above calculation of the network loss value can specifically be: calculating the cross-entropy loss value of the human segmentation network according to the cross-entropy loss function, and identifying the cross-entropy loss value as the network loss value; the cross-entropy loss function is as follows:
[0131] y ij =labels(t ij );
[0132] p ij =sigmoid(logout ij );
[0133] loss ij =-[y ij ×lnp ij +(1-y ij )×ln(1-p ij )];
[0134]
[0135] where t ij is the pixel value of the pixel at the i-th row and j-th column of the ground truth mask image; y ij is the probability value that the pixel at the i-th row and j-th column of the ground truth mask image is a human pixel; out ij is the pixel value of the pixel at the i-th row and j-th column of the predicted mask image; p ij is the probability value that the pixel at the i-th row and j-th column of the predicted mask image is a human pixel; loss ij is the deviation value corresponding to the pixel at the i-th row and j-th column of the predicted mask image; H and W are respectively the height and width of the predicted mask image (also representing the image size); loss _ c is the cross-entropy loss value.
[0136] In another possible implementation manner, calculating the network loss value may specifically be: calculating the alpha loss value of the human portrait segmentation network according to the alpha loss function, and identifying the alpha loss value as the network loss value; the alpha loss function is as follows:
[0137] y ij =labels(t ij )
[0138] p ij =sigmoid(logout ij )
[0139]
[0140] where t ij is the pixel value of the pixel at the i-th row and j-th column of the mask ground truth image; y ij is the probability value that the pixel at the i-th row and j-th column of the mask ground truth image is a human portrait pixel; out ij is the pixel value of the pixel at the i-th row and j-th column of the predicted mask image; p ij is the probability value that the pixel at the i-th row and j-th column of the predicted mask image is a human portrait pixel; H and W are respectively the height and width of the predicted mask image (which also represent the image size); loss _ α is the alpha loss value.
[0141] It should be understood that the value ranges of y ij and p ij are [0, 1]. Specifically, if when configuring the mask ground truth image, each pixel value within the marked human portrait area is set to 255 and each pixel value within the marked background area is set to 0, then the labels function is specifically y ij =t ij / 255; if when configuring the mask ground truth image, each pixel value within the marked human portrait area is set to 0 and each pixel value within the marked background area is set to 255, then the labels function is specifically y ij =(255 - t ij ) / 255. The sigmoid function can refer to the relevant descriptions in the prior art and will not be elaborated here.
[0142] In S10-e, adjust the weight parameters of the human portrait segmentation network according to the network loss value.
[0143] In this embodiment, during the process of training the human portrait segmentation network, each training image corresponds to a network loss value. Taking one of the training images as an example for illustration, this training image corresponds to the network loss value, and the network loss value contains the output correction information for the human portrait segmentation network. The weight parameters of the human portrait segmentation network are adjusted multiple times according to the network loss values corresponding to multiple training images, so that the human portrait segmentation network after adjusting the weight parameters can be used for subsequent human portrait segmentation to improve the accuracy of human portrait segmentation.
[0144] In this embodiment, the human portrait segmentation network is trained through this training image set. Specifically, the weight parameters of the human portrait segmentation network are adjusted by using the network loss value to improve the output accuracy of the human portrait segmentation network, so as to improve the accuracy of subsequent human portrait segmentation.
[0145] Figure 11 The flowchart of the method provided in the seventh embodiment of the present application is shown. Refer to Figure 11 , relative to the above Figure 10 described embodiment, the method 10-d provided in this embodiment includes S10-d1 to S10-d4, which are specifically described in detail as follows:
[0146] Furthermore, calculating the network loss value of the human portrait segmentation network according to the predicted mask image and the mask ground truth image includes:
[0147] In this embodiment, the network loss value includes an edge loss value, and the edge loss value is used to describe the deviation between the edge information of the human portrait region in the predicted mask image and the mask ground truth image.
[0148] In S10-d1, a predicted probability image corresponding to the predicted mask image is obtained according to the softmax function.
[0149] In this embodiment, the softmax function is specifically: y_p ij = softmax(out ij ); where y_p ij is the pixel value of the pixel at the i-th row and j-th column of the predicted probability image, and is also the probability value that the pixel at the i-th row and j-th column of the predicted mask image is a human portrait pixel; out ij is the pixel value of the pixel at the i-th row and j-th column of the predicted mask image; the softmax function can refer to the relevant descriptions in the prior art and will not be elaborated here. It should be understood that the value range of y_p ij is [0, 1].
[0150] In S10-d2, the mask ground truth image is dilated and eroded to obtain a dilated ground truth image and an eroded ground truth image corresponding to the mask ground truth image.
[0151] In this embodiment, in order to calculate the edge loss value, it is necessary to extract the edge information in the masked ground truth image. The above-mentioned dilation and erosion processing of the masked ground truth image is performed to obtain a dilated ground truth image and an eroded ground truth image corresponding to the masked ground truth image. Preferably, specifically, it can be: using the standard dilation and erosion operations of tensorflow with a kernel parameter of [5, 5, 2] to perform the standard dilation processing on the masked ground truth image to obtain the dilated ground truth image, and performing the erosion processing on the masked ground truth image to obtain the eroded ground truth image. The specific implementation can refer to the relevant descriptions in the prior art and will not be elaborated here.
[0152] In S10-d3, a boundary ground truth image is generated according to the dilated ground truth image and the eroded ground truth image.
[0153] In this embodiment, the above-mentioned generation of the boundary ground truth image according to the dilated ground truth image and the eroded ground truth image can specifically be: performing a subtraction operation on the dilated ground truth image and the eroded ground truth image to obtain the boundary ground truth image, that is, the boundary ground truth image is the result of the morphological gradient processing (subtracting the dilation result from the erosion result) of the masked ground truth image. For details, please refer to Figure 12 , Figure 12 which shows a schematic diagram of the masked ground truth image and the boundary ground truth image provided in this embodiment. Figure 12 In the shown masked ground truth image, the white part is the portrait area. Figure 12 In the shown boundary ground truth image, the white part is the edge of the portrait area.
[0154] In S10-d4, each pixel value included in the predicted probability image, each pixel value included in the masked ground truth image, and each pixel value included in the boundary ground truth image are input into a preset edge loss function for calculation to obtain the edge loss value of the portrait segmentation network.
[0155] In this embodiment, the edge loss function is as follows:
[0156] loss _ edge = ∑{-edge_t ij *[y_t ij *ln(y_p ij ) + (1 - y_t ij )*ln(1 - y_p ij )]}
[0157] where the loss_edge is the edge loss value; the edge_t ij is the pixel value at the i-th row and j-th column in the boundary ground truth image, that is, the probability value that the pixel at the i-th row and j-th column in the masked ground truth image is the edge of the portrait area; the y_t ijis the probability value that the pixel at the i-th row and j-th column in the true value image of the mask is a portrait pixel; this y_p ij is the probability value that the pixel at the i-th row and j-th column in the predicted probability image is a portrait pixel. It should be understood that this y_p ij , y_t ij and edge_t ij have a value range of [0, 1]; this y_t ij The calculation of can refer to the calculation of this y in S10-d above ij and will not be elaborated here.
[0158] It should be understood that in this embodiment, the network loss value includes the edge loss value: Exemplarily, the network loss value is the sum of the edge loss value and the above-mentioned cross-entropy loss value; Exemplarily, the network loss value is the sum of the edge loss value and the above-mentioned alpha loss value. Preferably, the network loss value is the sum of the edge loss value, the above-mentioned cross-entropy loss value, and the edge loss values before each decoding unit; Refer to Figure 7 , the edge loss values before each decoding unit are also the loss values of the M-th upsampled feature image. Denote the edge loss value of the M-th upsampled feature image as the M-th edge loss value. The M-th edge loss value is used to describe the edge loss of the part of the portrait segmentation network before the (M - 1)-th decoding unit. The calculation of the M-th edge loss can refer to the calculation of the edge loss value in this embodiment. It should be noted that when calculating the M-th edge loss, use the training image as the input, perform a channel adjustment convolution process and a point convolution process on the M-th upsampled feature image corresponding to the training image to adjust the number of channels and the size of the M-th upsampled feature image corresponding to the training image, and obtain the M-th loss image with the same number of channels and size as the portrait mask image; Replace the predicted mask image with the M-th loss image corresponding to the training image, and refer to the calculation steps of the edge loss value in this embodiment to calculate the M-th edge loss.
[0159] In this embodiment, by designing the edge loss function, when adjusting the weight parameters of the portrait segmentation network according to the network loss value, the network loss value includes the loss of the edge of the portrait area in the predicted mask image (the above-mentioned edge loss value), which can improve the effect of adjusting the weight parameters of the portrait segmentation network and further improve the output accuracy of the portrait segmentation network.
[0160] Figure 13 FIG. shows a schematic diagram of an application scenario provided by an embodiment of the present application, which is described in detail as follows:
[0161] See Figure 13 , compare Figure 3 , Figure 13The input size of the human portrait segmentation network is 304*304, and the number of channels is 3. That is, the size of the original human portrait image is 304*304, and the number of channels is 3, representing the three RGB channels; the size of the human portrait mask image is 304*304, and the number of channels is 2, representing the black and white channels. In one channel, black represents the human portrait area and white represents the background area. In the other channel, black represents the background area and white represents the human portrait area. The network loss value is the sum of the above-mentioned edge loss value, the above-mentioned cross-entropy loss value, and the edge loss values of each of the shown loss images, that is Figure 13 the sum of the edge loss values of the final output of the shown human portrait segmentation network, the first loss image, the second loss image, the third loss image, and the fourth loss image, and the cross-entropy loss value of the final output of the human portrait segmentation network. After training the human portrait segmentation network according to the network loss, test pictures are collected to test the human portrait segmentation network. When testing, the mean intersection over union (mIoU) of the predicted mask image and the ground truth mask image is specifically 0.94946, which can represent the effect of the human portrait segmentation network. The formula for calculating the mean intersection over union mIoU of the predicted mask image and the ground truth mask image is:
[0162]
[0163] where TP is the true positive, that is, the number of pixels where the human portrait pixels in the predicted mask image are in the same position as the human portrait pixels in the ground truth mask image; FP is the false positive, that is, the number of pixels where the human portrait pixels in the predicted mask image are in the same position as the background pixels in the ground truth mask image; FN is the false negative, that is, the number of pixels where the background pixels in the predicted mask image are in the same position as the human portrait pixels in the ground truth mask image; n is the number of test pictures; MIOU is the average of the intersection over union ratios of each test picture, that is, the above-mentioned mean intersection over union.
[0164] Corresponding to the method described in the above embodiments, Figure 14 FIG. shows a schematic structural diagram of a device provided by an embodiment of the present application. For the sake of convenience of description, only the parts related to the embodiments of the present application are shown.
[0165] Referring to Figure 14 , the human portrait segmentation device includes:
[0166] A human portrait original image acquisition module, configured to acquire a human portrait original image, and the picture of the human portrait original image includes a human portrait;
[0167] A human portrait segmentation network module, configured to input the human portrait original image into a human portrait segmentation network for processing and output a human portrait mask image; wherein, the human portrait segmentation network includes an encoding module, a pyramid pooling module, and a decoding module;
[0168] A human portrait segmentation module, configured to divide the human portrait original image into a human portrait image and a background image according to the human portrait mask image.
[0169] It should be noted that for the information interaction, execution process, etc. between the above-mentioned devices, since they are based on the same concept as the method embodiments of this application, for their specific functions and the technical effects brought, reference can be specifically made to the method embodiment part, and details will not be elaborated here.
[0170] Those skilled in the art can clearly understand that for the convenience and conciseness of description, only the above-mentioned division of each functional unit and module is used as an example. In actual applications, the above functions can be allocated to different functional units and modules according to needs, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above. Each functional unit and module in the embodiment can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above integrated units can be implemented in the form of hardware or in the form of software functional units. In addition, the specific names of each functional unit and module are only for the convenience of mutual distinction and do not limit the protection scope of this application. The specific working processes of the units and modules in the above system can refer to the corresponding processes in the foregoing method embodiments and will not be elaborated here.
[0171] Figure 15 The structural schematic diagram of a terminal device provided by an embodiment of this application is shown. As Figure 15 shown, the terminal device 15 in this embodiment includes: at least one processor 150 ( Figure 15 only one is shown in the figure), a processor, a memory 151, and a computer program 152 stored in the memory 151 and executable on the at least one processor 150. When the processor 150 executes the computer program 152, the steps in any of the foregoing method embodiments are implemented.
[0172] The terminal device 15 may be a computing device such as a desktop computer, a notebook, a palm computer, and a cloud server. The terminal device 15 is preferably a mobile terminal. The terminal device may include, but is not limited to, a processor 150 and a memory 151. Those skilled in the art can understand that Figure 15 merely examples of the terminal device 15 do not constitute a limitation on the terminal device 15, and it may include more or fewer components than shown in the figure, or combine certain components, or different components. For example, it may also include input and output devices, network access devices, etc.
[0173] The so-called processor 150 may be a Central Processing Unit (CPU), and the processor 150 may also be other general-purpose processors, Digital Signal Processors (DSPs), Application Specific Integrated Circuits (ASICs), Field-Programmable Gate Arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.
[0174] In some embodiments, the memory 151 may be an internal storage unit of the terminal device 15, such as the hard disk or memory of the terminal device 15. In some other embodiments, the memory 151 may also be an external storage device of the terminal device 15, such as a plug-in hard disk, a Smart Media Card (SMC), a Secure Digital (SD) card, a Flash Card, etc., equipped on the terminal device 15. Further, the memory 151 may also include both the internal storage unit and the external storage device of the terminal device 15. The memory 151 is used to store an operating system, application programs, a BootLoader, data, and other programs, such as the program code of the computer program. The memory 151 may also be used to temporarily store data that has been output or is to be output.
[0175] An embodiment of the present application also provides a computer-readable storage medium storing a computer program, and when the computer program is executed by a processor, the steps in the above method embodiments can be implemented.
[0176] An embodiment of the present application provides a computer program product, and when the computer program product runs on a mobile terminal, the mobile terminal can implement the steps in the above method embodiments when executed.
[0177] When the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, to implement all or part of the processes in the above-described embodiment methods of this application, a computer program can be used to instruct relevant hardware to complete. This computer program can be stored in a computer-readable storage medium. When this computer program is executed by a processor, the steps of the above-described method embodiments can be implemented. Among them, this computer program includes computer program code, and this computer program code can be in the form of source code, object code, executable file, or some intermediate form, etc. This computer-readable medium can at least include: any entity or device that can carry the computer program code to the photographing device / terminal device, recording medium, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signal, telecommunication signal, and software distribution medium. For example, a USB flash drive, a mobile hard disk, a magnetic disk, or an optical disc, etc. In some jurisdictions, according to legislation and patent practice, the computer-readable medium cannot be an electrical carrier signal and a telecommunication signal.
[0178] In the above embodiments, the descriptions of each embodiment have their own emphases. For the parts not detailed or recorded in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.
[0179] Those of ordinary skill in the art can realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods for each specific application to implement the described functions, but such implementation should not be considered to exceed the scope of this application.
[0180] In the embodiments provided in this application, it should be understood that the disclosed device / terminal device and method can be implemented in other ways. For example, the device / terminal device embodiments described above are only illustrative. For example, the division of the module or unit is only a logical function division. In actual implementation, there can be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection to each other can be through some interfaces. The indirect coupling or communication connection of the device or unit can be in an electrical, mechanical, or other form.
[0181] The units described above as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place, or they may be distributed across multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0182] The embodiments described above are only used to illustrate the technical solutions of the present application, rather than to limit it; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present application, and should all be included in the protection scope of the present application.
Claims
1. A method for human portrait segmentation, characterized in that, Including: Obtain an original portrait image, where the picture of the original portrait image contains a portrait; Input the original portrait image into a portrait segmentation network for processing, and output a portrait mask image; wherein, the portrait segmentation network includes an encoding module, a pyramid pooling module, and a decoding module; the encoding module uses depthwise separable convolutional layers to extract the feature information of the original portrait image; the decoding module includes multiple decoding units, and each decoding unit includes an upsampling layer, a feature fusion layer, and a conversion unit, and the conversion unit includes at least one depthwise separable convolutional layer, at least one channel adjustment convolutional layer, and at least one feature fusion block; the depthwise separable convolutional layer and the channel adjustment convolutional layer in the conversion unit are arranged in parallel, and the outputs of both are fused through the feature fusion block to obtain a feature image including the internal connections of each pixel point; Divide the original portrait image into a portrait image and a background image according to the portrait mask image.
2. The method according to claim 1, wherein The step of inputting the original portrait image into a portrait segmentation network for processing and outputting a portrait mask image includes: Input the original portrait image into the encoding module, and perform multiple convolutional processes on the original portrait image using the depthwise separable convolutional algorithm corresponding to the encoding module, and output N downsampled feature images, where N is an integer greater than 2; Input the Nth downsampled feature image into the pyramid pooling module for processing, and output the Nth upsampled feature image; Input the Nth upsampled feature image and the N - 1 downsampled feature images other than the Nth downsampled feature image among the N downsampled feature images into the decoding module for processing, and output a first upsampled feature image; wherein, the decoding module includes N - 1 decoding units, and the (M - 1)th decoding unit is used to determine the (M - 1)th upsampled feature image according to the Mth upsampled feature image and the (M - 1)th downsampled feature image, where M is an integer less than or equal to N and M is greater than 0; Perform image masking processing on the first upsampled feature image to obtain a portrait mask image.
3. The method according to claim 2, wherein The step of inputting the Nth downsampled feature image into the pyramid pooling module for processing and outputting the Nth upsampled feature image includes: Respectively perform atrous convolution processing on the Nth downsampled feature image using a plurality of preset convolution kernels with different scales to obtain a plurality of atrous convolution feature images, and the plurality of atrous convolution feature images correspond one-to-one to the plurality of convolution kernels with different scales; Respectively perform convolution processing on each atrous convolution feature image to obtain a plurality of to-be-fused feature images corresponding one-to-one to the plurality of atrous convolution feature images; Fuse the plurality of to-be-fused feature images to obtain a pooled feature image; Perform channel splicing on the pooled feature image and the Nth downsampled feature image to obtain the Nth upsampled feature image.
4. The method according to claim 3, wherein The step of determining the (M - 1)th upsampled feature image according to the Mth upsampled feature image and the (M - 1)th downsampled feature image includes: Input the M-th upsampled feature image into the upsampling layer of the (M - 1)-th decoding unit, perform upsampled feature mapping on the M-th upsampled feature image, and output the (M - 1)-th initial feature image; Input the (M - 1)-th initial feature image and the (M - 1)-th downsampled feature image into the feature fusion layer of the (M - 1)-th decoding unit, perform feature fusion on the (M - 1)-th initial feature image and the (M - 1)-th downsampled feature image, and obtain the (M - 1)-th upsampled feature fusion image; Input the (M - 1)-th upsampled feature fusion image into the conversion unit of the (M - 1)-th decoding unit for processing, and output the (M - 1)-th upsampled feature image.
5. The method according to claim 4, wherein The step of inputting the (M - 1)-th upsampled feature fusion image into the conversion unit of the (M - 1)-th decoding unit for processing and outputting the (M - 1)-th upsampled feature image includes: Input the (M - 1)-th upsampled feature fusion image into at least one depthwise separable convolutional layer for processing, and output the (M - 1)-th image to be converted; Input the (M - 1)-th upsampled feature fusion image into at least one channel adjustment convolutional layer for processing, and output the (M - 1)-th channel adjustment image; Input the (M - 1)-th image to be converted and the (M - 1)-th channel adjustment image into at least one feature fusion block for feature fusion, and output the (M - 1)-th upsampled feature image.
6. The method according to any one of claims 1-5, characterized in that, Before inputting the original portrait image into the portrait segmentation network for processing and outputting the portrait mask image, the method further includes: Obtain a training image set, the training image set includes a number of training images, and the scenes of the training images contain portraits; the training images are configured with associated mask ground truth images; Input the number of training images into the portrait segmentation network for processing, and output prediction mask images corresponding to the number of training images one by one; Calculate the network loss value of the portrait segmentation network according to the prediction mask image and the mask ground truth image, the prediction mask image is any one of the number of prediction mask images, and the mask ground truth image is any one of the number of mask ground truth images; Adjust the weight parameters of the portrait segmentation network according to the network loss value.
7. The method according to claim 6, wherein The network loss value includes an edge loss value, and the step of calculating the network loss value of the portrait segmentation network according to the prediction mask image and the mask ground truth image includes: Obtain the prediction probability image corresponding to the prediction mask image according to the softmax function; Perform dilation and erosion processing on the mask ground truth image to obtain the dilated ground truth image and the eroded ground truth image corresponding to the mask ground truth image; Generate a boundary ground truth image according to the dilated ground truth image and the eroded ground truth image; Input each pixel value included in the prediction probability image, each pixel value included in the mask ground truth image, and each pixel value included in the boundary ground truth image into a preset edge loss function for calculation, and obtain the edge loss value of the portrait segmentation network; Wherein, the edge loss function is: loss _ edge = ∑{-edge_t ij *[y_t ij *ln(y_p ij ) + (1 - y_t ij )*ln(1 - y_p ij )]} The loss_edge is the edge loss value, and the edge_t ij is the pixel value at the i-th row and j-th column in the ground truth boundary image; the y_t ij is the pixel value at the i-th row and j-th column in the ground truth mask image; the y_p ij is the pixel value at the i-th row and j-th column in the predicted probability image.
8. A portrait segmentation device, characterized in that, Includes: A portrait original image acquisition module, which is used to acquire a portrait original image, and the picture of the portrait original image contains a portrait; A portrait segmentation network module, which is used to input the portrait original image into a portrait segmentation network for processing and output a portrait mask image; wherein, the portrait segmentation network includes an encoding module, a pyramid pooling module and a decoding module; the encoding module uses depthwise separable convolutional layers to replace standard convolutional layers; the decoding module includes a plurality of decoding units, and each decoding unit includes an upsampling layer, a feature fusion layer and a conversion unit, and the conversion unit includes at least one depthwise separable convolutional layer, at least one channel adjustment convolutional layer and at least one feature fusion block; the depthwise separable convolutional layer and the channel adjustment convolutional layer in the conversion unit are arranged in parallel, and the outputs of the two are fused through the feature fusion block to obtain a feature image including the internal connection of each pixel point; A portrait segmentation module, which is used to divide the portrait original image into a portrait image and a background image according to the portrait mask image.
9. A terminal device, characterized in that, The terminal device includes a memory, a processor, and a computer program stored in the memory and executable on the processor, and when the computer program is executed by the processor, the method described in any one of claims 1 to 7 is implemented.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the method described in any one of claims 1 to 7 is implemented.
Citation Information
Patent Citations
Portrait segmentation method and device
CN109493350A
Image processing method, device and equipment and storage medium
CN111369581A