Video Portrait Segmentation Method, Device, Equipment and Medium Based on Attention Mechanism

Through the video portrait segmentation method based on attention mechanism, combined with shallow and deep feature processing and edge softening technology, the problems of low segmentation accuracy and edge instability in the existing methods are solved, and higher quality video portrait segmentation is achieved.

CN115171186BActive Publication Date: 2025-07-29SHENZHEN WONDERSHARE SOFTWARE CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210772408.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-06-30
Publication Date
2025-07-29
Estimated Expiration
2042-06-30

AI Technical Summary

Technical Problem

The existing video portrait segmentation method cannot effectively consider global context information, resulting in low segmentation accuracy and serious edge jagging and instability.

Method used

The video portrait segmentation method based on attention mechanism is adopted, shallow and deep features are extracted through a pre-trained segmentation model, feature processing and upsampling are performed in combination with the attention module, and then feature addition and segmentation are performed. Finally, the segmentation accuracy is improved through mean filtering and edge softening treatment.

Benefits of technology

Improve the integrity and accuracy of video portrait segmentation, reduce edge jagging and instability, and improve segmentation accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115171186B_ABST
    Figure CN115171186B_ABST
Patent Text Reader

Abstract

The present application relates to a video human portrait segmentation method, device, equipment and medium based on an attention mechanism. The method includes obtaining a human body image in a video as a human body image to be segmented; extracting features of the human body image to be segmented through a pre-trained segmentation model to obtain shallow features and first deep features; performing feature processing on the shallow features and the first deep features to obtain an attention image and generating second deep features; performing upsampling processing on the second deep features to obtain first sampled features, and performing feature addition processing on the first sampled features and the shallow features to obtain third deep features; performing feature segmentation processing on the third deep features to obtain an initial segmented human body image; performing mean filtering processing on the initial segmented human body image to obtain a filtered image, and performing edge softening processing based on the filtered image to obtain a target segmented human body image. The embodiments of the present invention are beneficial to providing the segmentation accuracy of video human portraits.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of semantic segmentation, and in particular, to a video portrait segmentation method, device, equipment and medium based on an attention mechanism. Background Art

[0002] Semantic segmentation is a very important direction in computer vision. Different from object detection and recognition, semantic segmentation realizes pixel-level classification of images. That is, a category is assigned to each pixel. Therefore, it can divide an image or video into multiple different blocks according to the similarities and differences of categories, so as to achieve the purpose of image semantic segmentation. Portrait segmentation belongs to a type of semantic segmentation. In an image or video frame, the human body is regarded as the foreground category, and the rest are regarded as the background category, dividing the entire picture into two categories. This technology currently has a wide range of applications, and has in-depth and extensive applications in entertainment scenarios such as portrait special effects, background replacement in video conferences or live broadcasts, and so on.

[0003] Existing video portrait segmentation methods mainly use deep learning neural networks. Through a convolutional neural network assumed artificially, a human body picture is mapped to a high-dimensional vector. Usually, in the training of the neural network model for feature extraction, it will drive the human body features after mapping to have as large a feature distance as possible between different people and as small as possible for the same person in its feature space, so as to output a human body segmentation image. However, this method cannot effectively consider global context information, and the edges of the segmented video portrait are prone to jagged and unstable phenomena, resulting in low accuracy of video portrait segmentation. Summary of the Invention

[0004] The purpose of the embodiments of this application is to provide a video portrait segmentation method, device, equipment and medium based on an attention mechanism to improve the segmentation accuracy of video portraits.

[0005] To solve the above technical problems, the embodiments of this application provide a video portrait segmentation method based on an attention mechanism, including:

[0006] Obtain a human body image in the video as the human body image to be segmented;

[0007] Extract features from the human body image to be segmented through a pre-trained segmentation model to obtain shallow features and first deep features;

[0008] Perform feature processing on the shallow features and the first deep features through a first attention module to obtain an attention image, and multiply the attention image and the deep features in matrix form to obtain second deep features;

[0009] Based on the shallow features, perform upsampling on the second deep features to obtain first sampled features, and perform feature addition on the first sampled features and the shallow features to obtain third deep features;

[0010] Perform feature segmentation on the third deep features through a second attention module to obtain an initial segmented human body image;

[0011] Perform mean filtering on the initial segmented human body image to obtain a filtered image, and perform edge softening on the filtered image to obtain a target segmented human body image.

[0012] To solve the above technical problems, an embodiment of the present application provides a video portrait segmentation device based on an attention mechanism, including:

[0013] A human body image to be segmented acquisition module, configured to acquire a human body image in a video as a human body image to be segmented;

[0014] A human body image feature extraction module, configured to extract features of the human body image to be segmented through a pre-trained segmentation model to obtain shallow features and first deep features;

[0015] A second deep feature generation module, configured to perform feature processing on the shallow features and the first deep features through a first attention module to obtain an attention image, and perform matrix multiplication on the attention image and the deep features to obtain second deep features;

[0016] A third deep feature generation module, configured to perform upsampling on the second deep features based on the shallow features to obtain first sampled features, and perform feature addition on the first sampled features and the shallow features to obtain third deep features;

[0017] An initial segmented human body image generation module, configured to perform feature segmentation on the third deep features through a second attention module to obtain an initial segmented human body image;

[0018] A target segmented human body image generation module, configured to perform mean filtering on the initial segmented human body image to obtain a filtered image, and perform edge softening on the filtered image to obtain a target segmented human body image.

[0019] To solve the above technical problems, a technical solution adopted by the present invention is: to provide a computer device, including one or more processors; a memory for storing one or more programs, so that one or more processors implement the video portrait segmentation method based on an attention mechanism described in any one of the above.

[0020] To solve the above technical problems, a technical solution adopted by the present invention is: a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the video portrait segmentation method based on the attention mechanism described in any one of the above is implemented.

[0021] The embodiments of the present invention provide a video portrait segmentation method, device, equipment and medium based on the attention mechanism. Among them, the method includes: obtaining a human body image in a video as a human body image to be segmented; extracting features of the human body image to be segmented through a pre-trained segmentation model to obtain shallow features and first deep features; performing feature processing on the shallow features and the first deep features through a first attention module to obtain an attention image, and multiplying the attention image with the deep features in matrix to obtain second deep features; based on the shallow features, performing upsampling processing on the second deep features to obtain a first sampled feature, and adding the first sampled feature and the shallow features in feature to obtain third deep features; performing feature segmentation processing on the third deep features through a second attention module to obtain an initial segmented human body image; performing mean filtering processing on the initial segmented human body image to obtain a filtered image, and performing edge softening processing based on the filtered image to obtain a target segmented human body image. The embodiments of the present invention extract the shallow features and deep features of the human body image in the video, combine the shallow features and deep features, thereby improving the integrity and accuracy of portrait segmentation. At the same time, combined with the attention mechanism, the details of portrait segmentation are further improved, and the edge softening processing is performed on the segmented human body image, reducing the serrations and instability of the edges of the video portrait segmentation, which is beneficial to providing the segmentation accuracy of the video portrait. Description of the Drawings

[0022] In order to more clearly illustrate the solutions in the present application, the following will briefly introduce the drawings required to be used in the description of the embodiments of the present application. Obviously, the following drawings are some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0023] Figure 1 It is a flowchart of an implementation of the video portrait segmentation method based on the attention mechanism provided by the embodiments of the present application;

[0024] Figure 2 It is another flowchart of an implementation of a sub-process in the video portrait segmentation method based on the attention mechanism provided by the embodiments of the present application;

[0025] Figure 3 It is another flowchart of an implementation of a sub-process in the video portrait segmentation method based on the attention mechanism provided by the embodiments of the present application;

[0026] Figure 4 It is another implementation flowchart of a sub - process in the video human - portrait segmentation method based on the attention mechanism provided by an embodiment of the present application;

[0027] Figure 5 It is another implementation flowchart of a sub - process in the video human - portrait segmentation method based on the attention mechanism provided by an embodiment of the present application;

[0028] Figure 6 It is another implementation flowchart of a sub - process in the video human - portrait segmentation method based on the attention mechanism provided by an embodiment of the present application;

[0029] Figure 7 It is another implementation flowchart of a sub - process in the video human - portrait segmentation method based on the attention mechanism provided by an embodiment of the present application;

[0030] Figure 8 It is a schematic diagram of a video human - portrait segmentation device based on the attention mechanism provided by an embodiment of the present application;

[0031] Figure 9 It is a schematic diagram of a computer device provided by an embodiment of the present application. Detailed implementation manners

[0032] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which this application belongs; the terms used in the specification of this application are only for the purpose of describing specific embodiments and are not intended to limit this application; the terms "including" and "having" and any variations thereof in the specification and claims of this application and the above - mentioned drawings are intended to cover non - exclusive inclusion. The terms "first", "second", etc. in the specification and claims of this application or the above - mentioned drawings are used to distinguish different objects and not to describe a specific order.

[0033] Referring to "embodiments" herein means that specific features, structures, or characteristics described in connection with the embodiments can be included in at least one embodiment of this application. The phrase appears in various places in the specification and does not necessarily refer to the same embodiment, nor is it an independent or alternative embodiment mutually exclusive with other embodiments. Those skilled in the art will explicitly and implicitly understand that the embodiments described herein can be combined with other embodiments.

[0034] In order to enable those skilled in the technical field to better understand the solution of this application, the technical solutions in the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings.

[0035] The present invention will be described in detail below with reference to the accompanying drawings and embodiments.

[0036] It should be noted that the video portrait segmentation method based on the attention mechanism provided by the embodiments of the present application is generally executed by a server. Correspondingly, the video portrait segmentation device based on the attention mechanism is generally configured in the server.

[0037] Please refer to Figure 1 , Figure 1 which shows a specific implementation of the video portrait segmentation method based on the attention mechanism.

[0038] It should be noted that if there are substantially the same results, the method of the present invention is not limited to Figure 1 the process sequence shown, and the method includes the following steps:

[0039] S1: Obtain a human body image in the video as the human body image to be segmented.

[0040] Specifically, obtain the video that needs portrait segmentation, and perform frame splitting on the video to obtain each frame of portrait image in the video, and use it as the segmented human body image. Further, a human body image can be directly obtained as the human body image to be segmented.

[0041] S2: Extract features from the human body image to be segmented through a pre-trained segmentation model to obtain shallow features and first deep features.

[0042] Specifically, the segmentation model in the embodiments of the present application uses mobilenetV2 as the backbone network, and preferentially selects 0.25 as the network width coefficient to reduce network parameters, thereby improving the portrait segmentation speed. Through training the segmentation model, a pre-trained segmentation model is obtained. The output of the 4th layer of the backbone network of the pre-trained segmentation model is used as shallow features, that is, the features of the 4-fold downsampling; the output of the 14th layer is used as the first deep features.

[0043] S3: Perform feature processing on the shallow features and the first deep features through a first attention module to obtain an attention image, and multiply the attention image with the deep features to obtain second deep features.

[0044] Specifically, perform feature processing on the shallow features and the first deep features through a first attention module (Attention Fusion Module) to generate an attention image with weight information, and then multiply the attention image with the deep features to obtain second deep features.

[0045] Please refer to Figure 2 , Figure 2 which shows a specific implementation of step S3, described in detail as follows:

[0046] S31: In the first attention module, based on the shallow features, upsample the first deep feature to obtain a second sampled feature.

[0047] Specifically, in the first attention module, based on the shallow features, upsample the first deep feature so that the first deep feature has the same resolution as the shallow features, thereby obtaining the second sampled feature.

[0048] S32: Concatenate the second sampled feature with the shallow features to obtain a concatenated feature.

[0049] Specifically, after performing convolution processing on the shallow features through a 3×3 convolutional layer, concatenate it with the second sampled feature to obtain a concatenated feature.

[0050] S33: Through matrix feature extraction processing on the concatenated feature, obtain a target matrix, and based on the target matrix, generate a shallow matrix feature corresponding to the shallow features and a first deep matrix feature corresponding to the first deep feature.

[0051] Specifically, after performing convolution, activation, pooling, etc. on the concatenated feature, thereby implementing matrix feature extraction processing on the concatenated feature to obtain a target matrix, and then based on the target matrix, generate a shallow matrix feature corresponding to the shallow features and a first deep matrix feature corresponding to the first deep feature.

[0052] Please refer to Figure 3 , Figure 3 which shows a specific implementation manner of step S33, described in detail as follows:

[0053] S331: Through convolution processing and activation processing on the concatenated feature, obtain a convolutional feature.

[0054] S332: According to the global pooling method, perform pooling processing on the convolutional feature to obtain a pooled feature, and perform convolution processing and activation processing on the pooled feature to obtain a target matrix.

[0055] S333: Perform convolution processing on the shallow features and the first deep feature respectively to obtain a shallow convolutional feature and a first deep convolutional feature.

[0056] S334: Based on the target matrix, generate a shallow target matrix, and perform multiplication processing on the shallow convolutional feature and the shallow target matrix to obtain a shallow matrix feature.

[0057] S335: Perform multiplication processing on the first deep convolutional feature and the target matrix to obtain a first deep matrix feature.

[0058] Specifically, after the splicing feature is convolved through a 1×1 convolutional layer and then activated through a rectified linear unit (ReLU function), a convolutional feature is obtained. Then, through global pooling of the convolutional feature, a pooled feature is obtained. After that, the pooled feature is convolved through a 1×1 convolutional layer and then activated through a Sigmoid function to obtain a target matrix. Then, return the shallow feature and the first deep feature are respectively convolved through a 3×3 convolutional layer to obtain a shallow convolutional feature and a first deep convolutional feature. Multiply the first deep convolutional feature by the target matrix to obtain a first deep matrix feature. Then subtract the target matrix from 1 to obtain a shallow target matrix, and multiply the shallow convolutional feature by the shallow target matrix to obtain a shallow matrix feature.

[0059] S34: Upsample the first deep matrix feature to obtain a third sampled feature, and add the third sampled feature to the shallow matrix feature to obtain a second deep feature.

[0060] Specifically, upsample the first deep matrix feature so that the resolution of the first deep matrix feature is the same as that of the shallow matrix feature to obtain a third sampled feature, and then add the third sampled feature to the shallow matrix feature to obtain a second deep feature.

[0061] S4: Based on the shallow feature, upsample the second deep feature to obtain a first sampled feature, and add the first sampled feature to the shallow feature to obtain a third deep feature.

[0062] Specifically, based on the shallow feature, upsample the second deep feature so that the resolution of the second deep feature is the same as that of the shallow feature to obtain a first sampled feature, and then add the first sampled feature to the shallow feature to obtain a third deep feature.

[0063] S5: Feature-segment the third deep feature through a second attention module to obtain an initial segmented human image.

[0064] Specifically, feature-segment the third deep feature through a second attention module (Strip Attention Module), take the human body in the human image as the foreground, and segment the human body foreground and background in the human image to obtain an initial segmented image.

[0065] Please refer to Figure 4 , Figure 4 which shows a specific implementation manner of step S5, described in detail as follows:

[0066] S51: In the second attention module, perform convolution processing on the third deep feature respectively to obtain a first convolution feature, a second convolution feature, and a third convolution feature.

[0067] S52: Perform segmentation processing on the second convolution feature and the third convolution feature respectively to obtain a second segmentation feature and a third segmentation feature.

[0068] S53: Generate a segmentation matrix based on the first convolution feature, the second segmentation feature, and the third segmentation feature, and perform matrix addition processing on the segmentation matrix and the third convolution feature to obtain an initial segmented human body image.

[0069] Specifically, in the second attention module, perform convolution processing on the third deep feature through a 3*3 convolution layer respectively to obtain a first convolution feature C*H*W, a second convolution feature C*H*W, and a third convolution feature C*H*W. Perform segmentation processing on the second convolution feature C*H*W and the third convolution feature C*H*W respectively, that is, perform striping processing on the second convolution feature C*H*W and the third convolution feature C*H*W respectively, so that the feature is segmented into data blocks of the same size, thereby obtaining a second segmentation feature C*W and a third segmentation feature C*W. Then perform reshape processing on the second segmentation feature C*W to make it an array with the same dimension as the first convolution feature C*H*W, and perform matrix multiplication processing on the reshaped second segmentation feature C*W and the first convolution feature C*H*W to obtain a multiplication matrix N*W. Perform reshape processing on the third convolution feature C*H*W to make it an array with the same dimension as the multiplication matrix N*W, and then perform transpose processing on it, that is, change the order of its rows and columns, to obtain a transposed feature W*C. Then perform matrix multiplication processing on the transposed feature W*C and the multiplication matrix N*W to obtain a segmentation matrix, and finally perform matrix addition processing on the segmentation matrix and the third convolution feature to achieve human portrait segmentation, thereby obtaining an initial segmented human body image.

[0070] S6: Perform mean filtering processing on the initial segmented human body image to obtain a filtered image, and perform edge softening processing based on the filtered image to obtain a target segmented human body image.

[0071] Specifically, the above steps have performed segmentation processing on the human body image in the video to obtain an initial segmented human body image. However, the segmentation edge of the initial segmented human body image may have serrations and instability phenomena. Therefore, the embodiment of the present application performs mean filtering processing on the initial segmented human body image to obtain a filtered image, and performs edge softening processing based on the filtered image to obtain a target segmented human body image, realizing the reduction of serrations and instability phenomena at the segmentation edge of the video human portrait, thereby improving the segmentation accuracy of the video human portrait.

[0072] Please refer toFigure 5 , Figure 5 shows a specific implementation of step S6, which is described in detail as follows:

[0073] S61: Perform mean filtering on the initial segmented human body image to obtain a filtered image.

[0074] S62: Enlarge the filtered image to obtain a filtered and enlarged image, and perform median processing on the filtered and enlarged image to obtain a median-processed human body image.

[0075] S63: Perform erosion and edge filtering on the median-processed human body image to obtain a target segmented human body image.

[0076] Specifically, perform 5*5 mean filtering on the initial segmented human body image to obtain a filtered image, and then enlarge the filtered image to expand it by a preset size. The preset size is set according to the actual situation and is not limited here. In a specific embodiment, the filtered image is enlarged to one-fourth of the size of the image to be segmented to obtain a filtered and enlarged image. Then, take a preset value to perform median processing on the filtered and enlarged image to obtain a median-processed human body image, where the preset value is set according to the actual situation and is not limited here. In a specific embodiment, the preset value is 51. Then, perform erosion processing on the median-processed human body image with a kernel size of 11, and perform edge filtering on it with a kernel size of 5*5, so as to achieve softening of the edges of the filtered image, reduce the sawtooth and instability phenomena at the edges of the video portrait segmentation, and thus improve the segmentation accuracy of the video portrait. Among them, erosion processing refers to reducing and thinning the bright areas or white parts in the image, so that the running result image is smaller than the bright area of the original image.

[0077] Please refer to Figure 6 , Figure 6 shows a specific implementation before step S1, which is described in detail as follows:

[0078] S1A: Obtain a sample human body image and a labeled image corresponding to the sample human body image, where the labeled image includes a set of portrait foreground pixels and a true class label.

[0079] Specifically, before implementing step S1, the embodiment of the present application will first train the segmentation model to obtain a pre-trained segmentation model. Obtain a sample human body image and a labeled image corresponding to the sample human body image. Each pixel in the sample human body image is manually labeled one by one to obtain a labeled image. Among them, the labeled image includes a set of portrait foreground pixels and a true class label.

[0080] S1B: Extract features from the sample human body image through the segmentation model to obtain shallow training features and first deep training features, and output the first feature information of the shallow training features through the first output end of the segmentation model.

[0081] Among them, the first feature information includes the current sample portrait foreground pixel set and the current sample portrait class label.

[0082] Specifically, the segmentation model uses mobilenetV2 as the backbone network, takes 0.25 as the network width coefficient. Take the output of the 4th layer of the backbone network as the shallow training features, that is, the features of the 4-fold downsampling; take the output of the 14th layer as the first deep training features. Output the first feature information of the shallow training features through the first output end (outputhead1) of the segmentation model. Among them, the first output end (outputhead1) is a ConvBNReLU operator composed of a series connection of CONV (convolution), batchnormalization (batch normalization), and relu (non-linear unit). Among them, the first feature information includes the current sample portrait foreground pixel set and the current sample portrait class label.

[0083] S1C: Based on the labeled image and the first feature information, generate a dice loss and a cross-entropy loss, and add the dice loss and the cross-entropy loss to obtain the first target loss.

[0084] Specifically, compare and analyze the labeled image and the first feature information to generate a dice loss (Diceloss) and a cross-entropy loss (pixelloss).

[0085] Please refer to Figure 7 , Figure 7 which shows a specific implementation manner of step S1C, described in detail as follows:

[0086] S1C1: Calculate and process the portrait foreground pixel set and the current sample portrait foreground pixel set through the first preset formula to obtain the dice loss.

[0087] S1C2: Calculate and process the current sample portrait class label and the current sample portrait class label through the second preset formula to obtain the cross-entropy loss.

[0088] S1C3: Add the dice loss and the cross-entropy loss to obtain the first target loss.

[0089] The first preset formula:

[0090]

[0091] Among them, DiceLoss is the diced loss, X is the set of portrait foreground pixels, and Y is the set of portrait foreground pixels of the current sample;

[0092] Second preset formula:

[0093]

[0094] Among them, pixel is the cross-entropy loss, y pred is the portrait class label of the current sample, y true The portrait class label of the current sample, and classes represents the total number of classes.

[0095] Specifically, since the diced loss (Diceloss) tends to focus on the accuracy of the segmentation area, and the cross-entropy loss (pixelloss) tends to focus on the classification accuracy of individual pixels, the combination of the two enables the overall segmentation of the network to take into account both the accuracy of the area and the accuracy of pixel class judgment.

[0096] S1D: Based on the first target loss, the first output end of the segmentation model is trained by backpropagation to obtain the trained first output end and the target shallow features.

[0097] Specifically, based on the first target loss, the weights of the model are adjusted by backpropagation, and the extraction of shallow training features and the segmentation processing of the first feature information are re-performed until the first target loss reaches a preset threshold, at which point the adjustment of the model weights is stopped to obtain the trained first output end. Then, the feature information output by the first output end at this time is used as the target shallow features. The preset threshold is set according to the actual situation and is not limited here.

[0098] S1E: Based on the target shallow features and the first deep training features, the second deep training features are generated, and the second feature information of the second deep training features is output through the second output end of the segmentation model.

[0099] Specifically, the target shallow features and the first deep training features are input into the first attention module to generate an attention training image, and the attention training image is multiplied by the first deep training features in matrix form to obtain the second deep training features output by the second output end of the segmentation model.

[0100] S1F: Based on the second feature information and the annotation image, the second output end of the segmentation model is trained to obtain the trained second output end and the target deep features.

[0101] Specifically, a second target loss value is generated based on the second feature information and the annotated image. For the process of generating the second target loss value, please refer to steps S1C1 - S1C3. To avoid repetition, it will not be elaborated here. Then, the second output end of the segmentation model is trained through the second target loss value to obtain the trained second output end and the target deep features. For the specific training process, please refer to step S1D. To avoid repetition, it will not be elaborated here.

[0102] S1G: Based on the target deep features and the shallow training features, generate third deep training features, and based on the third deep training features and the annotated image, train the third output end of the segmentation model to obtain the trained third output end. And based on the trained first output end, the trained second output end, and the trained third output end, generate a pre-trained segmentation model.

[0103] Specifically, based on the shallow training features, perform upsampling on the target deep features so that the target deep features and the shallow training features have the same resolution to obtain the sampled training features. Then, add the sampled training features and the shallow training features to obtain the added features. Next, input the added features into the second attention module (StripAttention Module) for segmentation processing to obtain the third deep training features output by the third output end of the segmentation model. Based on the third deep training features and the annotated image, generate a third target loss value. For the process of generating the third target loss value, please refer to steps S1C1 - S1C3. To avoid repetition, it will not be elaborated here. Then, the third output end of the segmentation model is trained through the third target loss value to obtain the trained third output end. For the specific training process, please refer to step S1D. To avoid repetition, it will not be elaborated here. Then, based on the trained first output end, the trained second output end, and the trained third output end, generate a pre-trained segmentation model. Among them, the second output end and the third output end have the same network structure as the first output end (outputhead1).

[0104] In this embodiment, a human body image in a video is obtained as a human body image to be segmented. The pre-trained segmentation model is used to extract features from the human body image to be segmented, obtaining shallow features and first deep features. The first attention module is used to perform feature processing on the shallow features and the first deep features to obtain an attention image, and the attention image is multiplied by the deep features in matrix form to obtain second deep features. Based on the shallow features, the second deep features are upsampled to obtain first sampled features, and the first sampled features and the shallow features are added together in terms of features to obtain third deep features. The second attention module is used to perform feature segmentation on the third deep features to obtain an initial segmented human body image. The initial segmented human body image is subjected to mean filtering to obtain a filtered image, and edge softening is performed based on the filtered image to obtain a target segmented human body image. In the embodiment of the present invention, by extracting the shallow features and deep features of the human body image in the video, combining the shallow features and deep features, the integrity and accuracy of human portrait segmentation are improved. At the same time, combined with the attention mechanism, the details of human portrait segmentation are further enhanced, and the edge of the segmented human body image is softened to reduce the sawtooth and instability phenomena at the edge of the video human portrait segmentation. Moreover, the dice loss and the cross-entropy loss are combined to train the segmentation, achieving both the accuracy of video human portrait classification and the integrity of the segmented area, thus facilitating the improvement of the segmentation accuracy of video human portraits.

[0105] Please refer to Figure 8 , as an implementation of the above Figure 1 shown method, an embodiment of a video human portrait segmentation device based on an attention mechanism is provided in this application. This device embodiment corresponds to Figure 1 the shown method embodiment and can be specifically applied to various electronic devices.

[0106] As Figure 8 shown, the video human portrait segmentation device based on an attention mechanism in this embodiment includes: a human body image to be segmented acquisition module 71, a human body image feature extraction module 72, a second deep feature generation module 73, a third deep feature generation module 74, an initial segmented human body image generation module 75, and a target segmented human body image generation module 76, where:

[0107] The human body image to be segmented acquisition module 71 is used to obtain a human body image in a video as a human body image to be segmented;

[0108] The human body image feature extraction module 72 is used to extract features from the human body image to be segmented through a pre-trained segmentation model, obtaining shallow features and first deep features;

[0109] The second deep feature generation module 73 is used to perform feature processing on the shallow features and the first deep features through the first attention module to obtain an attention image, and multiply the attention image with the deep features in matrix form to obtain the second deep features;

[0110] The third deep feature generation module 74 is used to perform upsampling processing on the second deep features based on the shallow features to obtain the first sampled features, and perform feature addition processing on the first sampled features and the shallow features to obtain the third deep features;

[0111] The initial segmented human body image generation module 75 is used to perform feature segmentation processing on the third deep features through the second attention module to obtain the initial segmented human body image;

[0112] The target segmented human body image generation module 76 is used to perform mean filtering processing on the initial segmented human body image to obtain a filtered image, and perform edge softening processing based on the filtered image to obtain the target segmented human body image.

[0113] Furthermore, the second deep feature generation module 73 includes:

[0114] The second sampled feature generation unit is used to perform upsampling processing on the first deep features based on the shallow features in the first attention module to obtain the second sampled features;

[0115] The feature splicing unit is used to splice the second sampled features and the shallow features to obtain the spliced features;

[0116] The target matrix generation unit is used to perform matrix feature extraction processing on the spliced features to obtain the target matrix, and generate the shallow matrix features corresponding to the shallow features and the first deep matrix features corresponding to the first deep features based on the target matrix;

[0117] The third sampled feature generation unit is used to perform upsampling processing on the first deep matrix features, add the third sampled features and the shallow matrix features to obtain the second deep features.

[0118] Furthermore, the target matrix generation unit includes:

[0119] The convolutional feature generation sub-unit is used to perform convolutional processing and activation processing on the spliced features to obtain the convolutional features;

[0120] The pooling processing sub-unit is used to perform pooling processing on the convolutional features according to the global pooling method to obtain the pooled features, and perform convolutional processing and activation processing on the pooled features to obtain the target matrix;

[0121] A convolution processing subunit, configured to perform convolution processing on the shallow features and the first deep features respectively to obtain shallow convolution features and first deep convolution features;

[0122] A matrix multiplication subunit, configured to generate a shallow target matrix based on the target matrix, and perform a multiplication process on the shallow convolution features and the shallow target matrix to obtain shallow matrix features;

[0123] A first deep matrix feature generation subunit, configured to perform a multiplication process on the first deep convolution features and the target matrix to obtain first deep matrix features.

[0124] Further, the initial segmented human body image generation module 75 includes:

[0125] A third deep feature convolution unit, configured to perform convolution processing on the third deep features respectively in the second attention module to obtain first convolution features, second convolution features, and third convolution features;

[0126] A transposed matrix generation unit, configured to perform segmentation processing on the second convolution features and the third convolution features respectively to obtain second segmentation features and third segmentation features;

[0127] A segmentation matrix generation unit, configured to generate a segmentation matrix based on the first convolution features, the second segmentation features, and the third segmentation features, and perform a matrix addition process on the segmentation matrix and the third convolution features to obtain an initial segmented human body image.

[0128] Further, the target segmented human body image generation module 76 includes:

[0129] A filtered image generation unit, configured to perform mean filtering processing on the initial segmented human body image to obtain a filtered image;

[0130] A median processing unit, configured to perform magnification processing on the filtered image to obtain a magnified filtered image, and perform median processing on the magnified filtered image to obtain a median-processed human body image;

[0131] An edge filtering unit, configured to perform erosion and edge filtering processing on the median-processed human body image to obtain a target segmented human body image.

[0132] Further, before the human body image to be segmented acquisition module 71, there is also included:

[0133] A sample human body image acquisition module, configured to acquire a sample human body image and an annotation image corresponding to the sample human body image, where the annotation image includes a portrait foreground pixel set and a true class label;

[0134] The sample human body image extraction module is used to extract features from the sample human body image through a segmentation model, obtain shallow training features and first deep training features, and output the first feature information of the shallow training features through the first output end of the segmentation model. Among them, the first feature information includes the current sample portrait foreground pixel set and the current sample portrait category label;

[0135] The loss calculation module is used to generate a dice loss and a cross-entropy loss based on the annotated image and the first feature information, and add the dice loss and the cross-entropy loss to obtain the first target loss;

[0136] The first training module is used to train the first output end of the segmentation model in a backpropagation manner based on the first target loss, and obtain the trained first output end and the target shallow features;

[0137] The second feature information generation module is used to generate second deep training features based on the target shallow features and the first deep training features, and output the second feature information of the second deep training features through the second output end of the segmentation model;

[0138] The second training module is used to train the second output end of the segmentation model based on the second feature information and the annotated image, and obtain the trained second output end and the target deep features;

[0139] The third training module is used to generate third deep training features based on the target deep features and the shallow training features, and train the third output end of the segmentation model based on the third deep training features and the annotated image to obtain the trained third output end, and generate a pre-trained segmentation model based on the trained first output end, the trained second output end, and the trained third output end.

[0140] Further, the loss calculation module includes:

[0141] The dice loss calculation unit is used to calculate and process the portrait foreground pixel set and the current sample portrait foreground pixel set through a first preset formula to obtain the dice loss;

[0142] The cross-entropy loss calculation unit is used to calculate and process the current sample portrait category label and the current sample portrait category label through a second preset formula to obtain the cross-entropy loss;

[0143] The first target loss unit is used to add the dice loss and the cross-entropy loss to obtain the first target loss.

[0144] The first preset formula:

[0145]

[0146] Among them, DiceLoss is the dice loss, X is the set of portrait foreground pixels, and Y is the set of portrait foreground pixels of the current sample;

[0147] Second preset formula:

[0148]

[0149] Among them, pixel is the cross-entropy loss, y pred is the portrait class label of the current sample, y true is the portrait class label of the current sample, and classes represents the total number of classes.

[0150] To solve the above technical problems, the embodiments of the present application also provide a computer device. For details, please refer to Figure 9 , Figure 9 which is the basic structural block diagram of the computer device in this embodiment.

[0151] The computer device 8 includes a memory 81, a processor 82, and a network interface 83 that are communicatively connected to each other through a system bus. It should be noted that only a computer device 8 with three components, namely a memory 81, a processor 82, and a network interface 83, is shown in the figure. However, it should be understood that it is not required to implement all the shown components, and more or fewer components can be alternatively implemented. Among them, those skilled in the art of the present technology can understand that a computer device here is a device that can automatically perform numerical calculations and / or information processing according to pre-set or stored instructions, and its hardware includes but is not limited to a microprocessor, an application specific integrated circuit (ASIC), a field-programmable gate array (FPGA), a digital signal processor (DSP), an embedded device, etc.

[0152] The computer device can be a desktop computer, a notebook, a palm computer, a cloud server, or other computing devices. The computer device can interact with the user through a keyboard, a mouse, a remote control, a touchpad, a voice control device, or other means.

[0153] The memory 81 includes at least one type of readable storage medium, which includes flash memory, hard disk, multimedia card, card-type memory (such as SD or DX memory, etc.), random access memory (RAM), static random access memory (SRAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), programmable read-only memory (PROM), magnetic memory, magnetic disk, optical disc, etc. In some embodiments, the memory 81 may be an internal storage unit of the computer device 8, such as the hard disk or memory of the computer device 8. In other embodiments, the memory 81 may also be an external storage device of the computer device 8, such as a plug-in hard disk, SmartMedia Card (SMC), Secure Digital (SD) card, Flash Card, etc. equipped on the computer device 8. Of course, the memory 81 may also include both the internal storage unit and the external storage device of the computer device 8. In this embodiment, the memory 81 is generally used to store the operating system and various application software installed on the computer device 8, such as the program code of the video human portrait segmentation method based on the attention mechanism. In addition, the memory 81 may also be used to temporarily store various types of data that have been output or will be output.

[0154] In some embodiments, the processor 82 may be a central processing unit (CPU), controller, microcontroller, microprocessor, or other data processing chip. The processor 82 is generally used to control the overall operation of the computer device 8. In this embodiment, the processor 82 is used to run the program code stored in the memory 81 or process data, such as running the program code of the above-mentioned video human portrait segmentation method based on the attention mechanism to implement various embodiments of the video human portrait segmentation method based on the attention mechanism.

[0155] The network interface 83 may include a wireless network interface or a wired network interface, and the network interface 83 is generally used to establish a communication connection between the computer device 8 and other electronic devices.

[0156] The present application also provides another implementation manner, that is, to provide a computer-readable storage medium storing a computer program, and the computer program can be executed by at least one processor to enable the at least one processor to execute the steps of a video human portrait segmentation method based on the attention mechanism as described above.

[0157] Through the description of the above embodiments, those skilled in the art can clearly understand that the above-described embodiment methods can be implemented by means of software plus a necessary general hardware platform. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions for causing a terminal device (which can be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods of the various embodiments of the present application.

[0158] Obviously, the above-described embodiments are only a part of the embodiments of the present application, rather than all of the embodiments. The accompanying drawings show the preferred embodiments of the present application, but do not limit the patent scope of the present application. The present application can be implemented in many different forms. On the contrary, the purpose of providing these embodiments is to make the understanding of the disclosed content of the present application more thorough and comprehensive. Although the present application has been described in detail with reference to the foregoing embodiments, for those skilled in the art, they can still modify the technical solutions recorded in the foregoing specific embodiments, or perform equivalent replacements on some of the technical features. Any equivalent structure directly or indirectly using the content of the specification and drawings of the present application in other related technical fields is similarly within the scope of the patent protection of the present application.

Claims

1. A video portrait segmentation method based on an attention mechanism, characterized in that Including: Obtaining a sample human body image and an annotation image corresponding to the sample human body image, wherein the annotation image includes a set of portrait foreground pixels and a true class label; Extracting features from the sample human body image through a segmentation model to obtain shallow training features and first deep training features, and outputting first feature information of the shallow training features through a first output end of the segmentation model, wherein the first feature information includes a current set of portrait foreground pixels and a current portrait class label of the sample; Generating a dice loss and a cross-entropy loss based on the annotation image and the first feature information, and adding the dice loss and the cross-entropy loss to obtain a first target loss; Training the first output end of the segmentation model by means of backpropagation based on the first target loss to obtain a trained first output end and target shallow features; Generating second deep training features based on the target shallow features and the first deep training features, and outputting second feature information of the second deep training features through a second output end of the segmentation model; Training the second output end of the segmentation model based on the second feature information and the annotation image to obtain a trained second output end and target deep features; Generating third deep training features based on the target deep features and the shallow training features, and training a third output end of the segmentation model based on the third deep training features and the annotation image to obtain a trained third output end, and generating a pre-trained segmentation model based on the trained first output end, the trained second output end, and the trained third output end; Obtaining a human body image in a video as a human body image to be segmented; Extracting features from the human body image to be segmented through the pre-trained segmentation model to obtain shallow features and first deep features; Performing feature processing on the shallow features and the first deep features through a first attention module to obtain an attention image, and multiplying the attention image and the first deep features matrix-wise to obtain second deep features; Performing upsampling processing on the second deep features based on the shallow features to obtain first sampled features, and adding the first sampled features and the shallow features feature-wise to obtain third deep features; Performing feature segmentation processing on the third deep features through a second attention module to obtain an initial segmented human body image; Performing mean filtering processing on the initial segmented human body image to obtain a filtered image, and performing edge softening processing based on the filtered image to obtain a target segmented human body image.

2. The video portrait segmentation method based on the attention mechanism according to claim 1 is characterized in that The step of performing feature processing on the shallow features and the first deep features through a first attention module to obtain an attention image, and multiplying the attention image and the first deep features matrix-wise to obtain second deep features includes: In the first attention module, performing upsampling processing on the first deep features based on the shallow features to obtain second sampled features; Splicing the second sampling feature with the shallow feature to obtain a spliced feature; Performing matrix feature extraction processing on the spliced features to obtain a target matrix, and generating shallow matrix features corresponding to the shallow features and first deep matrix features corresponding to the first deep features based on the target matrix; The first deep matrix features are upsampled to obtain third sampling features, and the third sampling features are added to the shallow matrix features to obtain the second deep features.

3. The method for video human portrait segmentation based on the attention mechanism according to claim 2, wherein, The step of performing matrix feature extraction processing on the spliced features to obtain a target matrix, and generating shallow matrix features corresponding to the shallow features and generating first deep matrix features corresponding to the first deep features based on the target matrix, includes: Performing convolution processing and activation processing on the splicing features to obtain convolution features; Performing pooling processing on the convolution features according to a global pooling method to obtain pooling features, and performing convolution processing and activation processing on the pooling features to obtain the target matrix; Convolutionally processing the shallow features and the first deep features to obtain shallow convolutional features and first deep convolutional features; Based on the target matrix, a shallow target matrix is generated, and the shallow convolution feature is multiplied by the shallow target matrix to obtain the shallow matrix feature; The first deep convolution feature is multiplied by the target matrix to obtain the first deep matrix feature.

4. The method for video portrait segmentation based on the attention mechanism according to claim 1, wherein The step of performing feature segmentation processing on the third deep feature by the second attention module to obtain an initial segmented human body image includes: In the second attention module, the third deep features are respectively convolved to obtain a first convolution feature, a second convolution feature, and a third convolution feature; Segmenting the second convolution feature and the third convolution feature respectively to obtain a second segmentation feature and a third segmentation feature; Based on the first convolution feature, the second segmentation feature, and the third segmentation feature, a segmentation matrix is generated, and matrix addition processing is performed on the segmentation matrix and the third convolution feature to obtain the initial segmented human body image.

5. The method for video human portrait segmentation based on the attention mechanism according to claim 1, characterized in that The performing mean filtering on the initial segmented human body image to obtain a filtered image, and performing edge softening processing based on the filtered image to obtain a target segmented human body image, including: Performing mean filtering on the initial segmented human body image to obtain the filtered image; Amplifying the filtered image to obtain a filtered amplified image, and medianizing the filtered amplified image to obtain a medianized human body image; The medianized human body image is subjected to corrosion and edge filtering processing to obtain the target segmented human body image.

6. The method for video portrait segmentation based on the attention mechanism according to claim 1, wherein The step of generating a dicing loss and a cross entropy loss based on the labeled image and the first feature information, and adding the dicing loss to the cross entropy loss to obtain a first target loss includes: Calculating the portrait foreground pixel set and the current sample portrait foreground pixel set using a first preset formula to obtain the dicing loss; Calculating the current sample portrait category label and the current sample portrait category label using a second preset formula to obtain the cross entropy loss; Adding the dicing loss to the cross entropy loss to obtain the first target loss; The first preset formula: ; Wherein, DiceLoss is the dicing loss, X is the portrait foreground pixel set, and Y is the portrait foreground pixel set of the current sample; The second preset formula: ; Among them, pixelloss is the cross-entropy loss, and y pred is the current sample portrait class label, and y true is the current sample portrait class label, and classes represents the total number of classes.

7. A video portrait segmentation device based on an attention mechanism, characterized in that include: A sample human body image acquisition module is used to acquire a sample human body image and an annotated image corresponding to the sample human body image, wherein the annotated image includes a set of portrait foreground pixels and a true category label; a sample human image extraction module, configured to extract features from the sample human image using a segmentation model to obtain shallow training features and first deep training features, and output first feature information of the shallow training features through a first output terminal of the segmentation model, wherein the first feature information includes a set of foreground pixels of a current sample portrait and a category label of the current sample portrait; a loss calculation module, configured to generate a dicing loss and a cross entropy loss based on the annotated image and the first feature information, and add the dicing loss and the cross entropy loss to obtain a first target loss; A first training module is configured to train the first output end of the segmentation model by back propagation based on the first target loss to obtain a trained first output end and target shallow features; A second feature information generating module is configured to generate a second deep training feature based on the target shallow feature and the first deep training feature, and output second feature information of the second deep training feature through a second output end of the segmentation model; A second training module is configured to train the second output end of the segmentation model based on the second feature information and the annotated image to obtain a trained second output end and target deep features; a third training module, configured to generate a third deep training feature based on the target deep feature and the shallow training feature, and train the third output end of the segmentation model based on the third deep training feature and the annotated image to obtain a trained third output end, and generate a pre-trained segmentation model based on the trained first output end, the trained second output end, and the trained third output end; A human body image acquisition module to be segmented is used to acquire a human body image in a video as the human body image to be segmented; A human body image feature extraction module, configured to extract features from the human body image to be segmented using the pre-trained segmentation model to obtain shallow features and first deep features; a second deep feature generation module, configured to perform feature processing on the shallow features and the first deep features through the first attention module to obtain an attention image, and perform matrix multiplication on the attention image and the first deep features to obtain a second deep feature; The third deep feature generation module is configured to perform upsampling processing on the second deep feature based on the shallow feature to obtain a first sampled feature, and perform feature addition processing on the first sampled feature and the shallow feature to obtain a third deep feature; The initial segmented human body image generation module is configured to perform feature segmentation processing on the third deep feature through the second attention module to obtain an initial segmented human body image; The target segmented human body image generation module is configured to perform mean filtering processing on the initial segmented human body image to obtain a filtered image, and perform edge softening processing on the filtered image to obtain a target segmented human body image.

8. A computer device, characterized in that, It includes a memory and a processor. A computer program is stored in the memory, and when the processor executes the computer program, it implements the attention mechanism-based video portrait segmentation method according to any one of claims 1 to 6.

9. A computer-readable storage medium, characterized in that, A computer program is stored on the computer-readable storage medium, and when the computer program is executed by the processor, it implements the attention mechanism-based video portrait segmentation method according to any one of claims 1 to 6.

Citation Information

Patent Citations

  • Image segmentation processing method and device, computer equipment and storage medium

    CN113538480A

  • Video target segmentation method and device, equipment and medium

    CN113763385A