A method, apparatus, computer device, and storage medium for portrait segmentation
By introducing attention fusion module and strip attention module into the EfficientNet network, feature fusion and stripping are performed, the robustness of portrait segmentation in the prior art under the background changes is solved, and a higher accuracy portrait segmentation effect is achieved.
Patent Information
- Application Number
- CN202210501867.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-05-09
- Publication Date
- 2025-06-24
- Estimated Expiration
- 2042-05-09
AI Technical Summary
The existing portrait segmentation algorithm is extremely robust under the situation of multiple changes in image background, and cannot complete complex portrait segmentation tasks.
The EfficientNet network is used to perform multiple dimension scaling processing, combining the attention fusion module and the strip attention module, feature fusion and stripping processing are performed, and the target object in the output image is finally segmented.
It improves the accuracy of portrait segmentation, enhances the robustness of the multiple changes in image background, and can more effectively complete complex portrait segmentation tasks.
Smart Images

Figure CN114897913B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image processing, and particularly to a method, device, computer device and storage medium for portrait segmentation. Background Art
[0002] With the increasing maturity of image fusion and digital processing technologies, higher requirements are also put forward for object replacement technologies such as faces and objects. How to make the replaced object look natural and indistinguishable from artificial replacement components. For example, in movies, higher precision and finer texture details are required for face replacement technology to ensure the authenticity of characters.
[0003] To address the above problems, researchers use deep neural networks for target object segmentation. For example, in the color-based color digital image target segmentation technology, artificial neural network methods are mainly used to segment the interested or uninterested pixels of the face image using the color space. Subsequently, a method based on a deep convolutional neural network is also proposed, that is, face alignment and segmentation are achieved through the multi-task joint learning algorithm of CNN. This algorithm realizes the collaborative fusion of cross-layer features through the designed residual module, improving the segmentation effect.
[0004] However, the above algorithms as a whole rely on low-level visual information and artificial auxiliary information for shallow semantic segmentation, lacking key model training and deep semantic information, resulting in extremely poor robustness of the above algorithms to the changing image background and being unable to complete complex target object segmentation tasks. Summary of the Invention
[0005] The purpose of the present invention is to provide a method, device, computer device and storage medium for portrait segmentation, aiming to solve the problem that the existing segmentation algorithms have extremely poor robustness to the changing image background and are unable to complete complex portrait segmentation tasks.
[0006] In a first aspect, an embodiment of the present invention provides a method for portrait segmentation, which includes:
[0007] Input the picture to be segmented into the EfficientNet network for multiple dimensionality reduction processes, and output the corresponding dimensionality-reduced pictures respectively;
[0008] Input the last two output dimensionality-reduced pictures into the attention fusion module for feature fusion to obtain a fused picture, where the last two output dimensionality-reduced pictures are 7×7×192 pictures and 7×7×320 pictures;
[0009] Input the fused picture into the strip attention module for striping processing to obtain an output image;
[0010] Segment the target object in the output image and output the target segmentation image.
[0011] In a second aspect, an embodiment of the present invention provides a human portrait segmentation device, which includes:
[0012] A dimension scaling unit for inputting the picture to be segmented into an EfficientNet network for multiple dimension scaling processes and respectively outputting corresponding downscaled pictures;
[0013] A feature fusion unit for inputting the last two output downscaled pictures into an attention fusion module for feature fusion to obtain a fused picture, where the last two output downscaled pictures are 7×7×192 pictures and 7×7×320 pictures;
[0014] A striping processing unit for inputting the fused picture into a strip attention module for striping processing to obtain an output image;
[0015] A segmentation unit for segmenting the target object in the output image and outputting a target segmentation image.
[0016] In a third aspect, an embodiment of the present invention further provides a computer device, which includes a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, the human portrait segmentation method described in the first aspect above is implemented.
[0017] In a fourth aspect, an embodiment of the present invention further provides a computer-readable storage medium, where the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the processor is caused to execute the human portrait segmentation method described in the first aspect above.
[0018] An embodiment of the present invention discloses a human portrait segmentation method, device, computer device, and storage medium. The method includes inputting the picture to be segmented into an EfficientNet network for multiple dimension scaling processes and respectively outputting corresponding downscaled pictures; inputting the last two output downscaled pictures into an attention fusion module for feature fusion to obtain a fused picture, where the last two output downscaled pictures are 7×7×192 pictures and 7×7×320 pictures; inputting the fused picture into a strip attention module for striping processing to obtain an output image; segmenting the target object in the output image and outputting a target segmentation image. By improving the network structure, the embodiment of the present invention has the advantage of improving the accuracy of human portrait segmentation. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0020] Figure 1 It is a schematic flowchart of the human portrait segmentation method provided by an embodiment of the present invention;
[0021] Figure 2 It is a schematic sub - flowchart of the human portrait segmentation method provided by an embodiment of the present invention;
[0022] Figure 3 It is another schematic sub - flowchart of the human portrait segmentation method provided by an embodiment of the present invention;
[0023] Figure 4 It is another schematic sub - flowchart of the human portrait segmentation method provided by an embodiment of the present invention;
[0024] Figure 5 It is a schematic block diagram of the human portrait segmentation device provided by an embodiment of the present invention;
[0025] Figure 6 It is a schematic block diagram of the computer device provided by an embodiment of the present invention. Detailed implementation manners
[0026] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0027] It should be understood that when used in this specification and the appended claims, the terms "comprises" and "comprising" indicate the presence of the described features, wholes, steps, operations, elements, and / or components, but do not exclude the presence or addition of one or more other features, wholes, steps, operations, elements, components, and / or their combinations.
[0028] It should also be understood that the terms used in this specification of the present invention are only for the purpose of describing specific embodiments and are not intended to limit the present invention. As used in this specification of the present invention and the appended claims, unless the context clearly indicates otherwise, the singular forms "a", "an" and "the" are intended to include the plural forms.
[0029] It should be further understood that the term "and / or" used in this specification of the present invention and the appended claims refers to any combination and all possible combinations of one or more of the associated listed items, and includes these combinations.
[0030] Please refer to Figure 1 , Figure 1 It is a schematic flowchart of the human portrait segmentation method provided by an embodiment of the present invention;
[0031] As Figure 1 shown, the method includes steps S101 to S104.
[0032] S101. Input the picture to be segmented into the EfficientNet network for multiple dimensional scaling processes, and respectively output the corresponding downscaled pictures;
[0033] In this step, the EfficientNet network includes 7 convolutional modules.
[0034] S102. Input the last two output downscaled pictures into the attention fusion module for feature fusion to obtain a fused picture, where the last two output downscaled pictures are 7×7×192 pictures and 7×7×320 pictures;
[0035] S103. Input the fused picture into the strip attention module for striping processing to obtain an output image;
[0036] S104. Segment the target object in the output image and output the target segmentation image.
[0037] In this embodiment, the AttaNet network framework is adopted, and the EfficientNet network is used as the core Backbone, which can balance the three dimensions of width, depth, and resolution, thereby improving the accuracy and speed of the AttaNet network.
[0038] Through the EfficientNet network, the picture to be segmented is sequentially subjected to dimensional scaling processing of 7 convolutional modules, and the corresponding downscaled pictures are respectively output. Then, through the attention fusion module (AFM), the last two downscaled pictures are subjected to feature fusion to obtain a fused picture. Then, through the strip attention module (SAM), the fused picture is subjected to striping processing to obtain an output image. Different objects are distinguished in the output image by different color features, and target segmentation is performed according to the features of the target object (i.e., portrait features), and then the target segmentation image can be output.
[0039] In one embodiment, as Figure 2 shown, step S101 includes:
[0040] S201. Input a 224×224×32 picture into the first convolutional module for convolutional processing, and output a 112×112×32 picture;
[0041] Among them, the number of convolutional layers of the first convolutional module is 1, and the convolutional kernel is 3×3; that is, the 224×224×32 picture is subjected to 1 convolutional processing in the first convolutional module and outputs a 112×112×32 picture.
[0042] S202. Input the 112×112×32 picture into the second convolution module for convolution processing, and output a 56×56×24 picture;
[0043] Among them, the number of convolution layers of the second convolution module is 3, and the convolution kernel is 3×3; that is, the 112×112×32 picture undergoes 3 times of convolution processing in the second convolution module, and finally outputs a 56×56×24 picture;
[0044] S203. Input the 56×56×24 picture into the third convolution module for convolution processing, and output a 28×28×40 picture;
[0045] Among them, the number of convolution layers of the third convolution module is 2, and the convolution kernel is 5×5; that is, the 56×56×24 picture undergoes 2 times of convolution processing in the third convolution module, and finally outputs a 28×28×40 picture;
[0046] S204. Input the 28×28×40 picture into the fourth convolution module for convolution processing, and output a 28×28×80 picture;
[0047] Among them, the number of convolution layers of the fourth convolution module is 3, and the convolution kernel is 3×3; that is, the 28×28×40 picture undergoes 3 times of convolution processing in the fourth convolution module, and finally outputs a 28×28×80 picture;
[0048] S205. Input the 28×28×80 picture into the fifth convolution module for convolution processing, and output a 7×7×192 picture;
[0049] Among them, the number of convolution layers of the fifth convolution module is 7, and the convolution kernel is 5×5; that is, the 28×28×80 picture undergoes 7 times of convolution processing in the fifth convolution module, and finally outputs a 7×7×192 picture;
[0050] S206. Input the 7×7×192 picture into the sixth convolution module for convolution processing, and output a 7×7×320 picture.
[0051] Among them, the number of convolution layers of the sixth convolution module is 1, and the convolution kernel is 3×3. That is, the 7×7×192 picture undergoes 1 time of convolution processing in the sixth convolution module, and finally outputs a 7×7×320 picture.
[0052] Through the dimensionality reduction processing of the 7 convolution modules in steps S201 - S206, the overall performance of the network is improved to a greater extent.
[0053] In an embodiment, as Figure 3 shown, step S102 includes:
[0054] S301. Perform convolution processing on the 7×7×192 picture to obtain the first picture;
[0055] In this step, a convolutional layer with a convolutional kernel of 3×3 is used for convolutional processing;
[0056] S302. Upsample the 7×7×192 image, then perform convolution, pooling, and activation processing, and assign weights to obtain a second image;
[0057] In this step, after upsampling, perform Concate function processing, convolution processing with a 1×1 convolutional kernel, Relu activation function processing, Global pool pooling processing, 1×1 convolution processing, and Sigmoid activation function processing in sequence, and then assign the weight α to obtain a second image;
[0058] S303. Perform a multiplication operation on the first image and the second image, and then perform upsampling processing to obtain a third image;
[0059] In this step, use the image multiplication operation Multply for the multiplication operation;
[0060] S304. Perform convolutional processing on the 7×7×320 image to obtain a fourth image;
[0061] In this step, a convolutional layer with a convolutional kernel of 3×3 is used for convolutional processing;
[0062] S305. Perform convolution, pooling, and activation processing on the fourth image, and assign weights to obtain a fifth image;
[0063] In this step, perform Concate function processing, convolution processing with a 1×1 convolutional kernel, Relu activation function processing, Global pool pooling processing, 1×1 convolution processing, and Sigmoid activation function processing on the fourth image in sequence, and then assign the weight 1-α to obtain a fifth image;
[0064] S306. Perform a multiplication operation on the fourth image and the fifth image, and then perform upsampling processing to obtain a sixth image;
[0065] S307. Perform feature fusion on the third image and the sixth image to obtain a fused image.
[0066] In this embodiment, through the attention fusion module, the 7×7×192 image and the 7×7×320 image are subjected to feature fusion according to the process of steps S301 - S307, achieving a further improvement in speed while ensuring accuracy. It should be noted that in this embodiment, the Global attention mechanism is introduced in the attention fusion module for weight assignment, and feature fusion is performed in combination with weights, greatly improving the feature accuracy of the fused image.
[0067] In one embodiment, as Figure 4 shown, step S103 includes:
[0068] S401. Perform three convolution operations on the fused image respectively to obtain a first convolution image, a second convolution image, and a third convolution image;
[0069] In this step, a convolution layer with a convolutional kernel of 3×3 can be used to perform three convolution operations on the fused image respectively to obtain a first convolution image, a second convolution image, and a third convolution image;
[0070] S402. Perform a dimension transpose process on the first convolution image to obtain a transposed image;
[0071] S403. Perform a striping process and a dimension arrangement process on the second convolution image, and then perform a batch process with the transposed image to obtain an attention matrix;
[0072] In this step, after the second convolution image undergoes a Striping striping process and a Reshape dimension arrangement process and then an Affinity batch process with the transposed image, an attention matrix of N×W is obtained, removing the H direction and greatly reducing the complexity of global context encoding in the vertical direction;
[0073] S404. Perform a striping process and a dimension transpose process on the third convolution image, multiply it with the attention matrix and then perform an arrangement process, and then add it to the fused image to obtain an output image.
[0074] In this embodiment, the fused image is processed by a strip-shaped attention module for striping and outputting an image. The regions of each object are accurately divided in the output image, facilitating subsequent precise segmentation.
[0075] In one embodiment, step S104 includes:
[0076] Obtain the regional edge information of each object in the output image, and use a semantic segmentation network to perform segmentation processing on the target object according to the regional edge information of the target object to obtain a target segmentation image.
[0077] In this embodiment, a semantic segmentation network is used. Based on the accurately divided regions of each object in the output image, the target object is segmented according to the regional edge information of the target object to be segmented, and a target segmentation image with high accuracy can be obtained.
[0078] In one embodiment, the portrait segmentation method further includes:
[0079] Optimize the portrait segmentation network model according to the following loss function:
[0080] Loss = Length + λ·Region;
[0081] Among them, Length represents the distance function, λ represents the hyperparameter, and Region represents the region function.
[0082] In this embodiment, when segmenting small targets, due to the obvious imbalance between the background and the foreground, the background is often much larger than the foreground, resulting in the loss function being dominated by the background and the segmentation effect being poor. Therefore, this embodiment introduces the cross-entropy loss function to improve the segmentation effect of small targets.
[0083] An embodiment of the present invention further provides a human portrait segmentation device, which is used to execute any embodiment of the foregoing human portrait segmentation method. Specifically, please refer to Figure 5 , Figure 5 which is a schematic block diagram of the human portrait segmentation device provided by the embodiment of the present invention.
[0084] As Figure 5 shown, the human portrait segmentation device 500 includes: a dimension scaling unit 501, a feature fusion unit 502, a striping processing unit 503, and a segmentation unit 504.
[0085] The dimension scaling unit 501 is used to input the picture to be segmented into the EfficientNet network for multiple dimension scaling processes, and respectively output the corresponding downscaled pictures;
[0086] The feature fusion unit 502 is used to input the last two output downscaled pictures into the attention fusion module for feature fusion to obtain a fused picture, where the last two output downscaled pictures are 7×7×192 pictures and 7×7×320 pictures;
[0087] The striping processing unit 503 is used to input the fused picture into the strip attention module for striping processing to obtain an output image;
[0088] The segmentation unit 504 is used to segment the target object in the output image and output the target segmentation image.
[0089] This device adopts the AttaNet network framework, uses the EfficientNet network as the core Backbone, balances the three dimensions of width, depth, and resolution, thereby improving the accuracy and speed of the AttaNet network; by improving the network structure, it has the advantage of improving the accuracy of human portrait segmentation.
[0090] Those skilled in the art can clearly understand that for the convenience and conciseness of description, the specific working processes of the above-described device and units can refer to the corresponding processes in the foregoing method embodiments, and will not be repeated here.
[0091] The above-mentioned portrait segmentation device can be implemented in the form of a computer program, which can run on a computer device as shown in Figure 6 the following.
[0092] Please refer to Figure 6 . Figure 6 FIG. is a schematic block diagram of a computer device provided by an embodiment of the present invention. The computer device 600 is a server, and the server can be an independent server or a server cluster composed of multiple servers.
[0093] Referring to Figure 6 , the computer device 600 includes a processor 602, a memory, and a network interface 605 connected through a system bus 601. Among them, the memory can include a non-volatile storage medium 603 and an internal memory 604.
[0094] The non-volatile storage medium 603 can store an operating system 6031 and a computer program 6032. When the computer program 6032 is executed, the processor 602 can be caused to execute the portrait segmentation method.
[0095] The processor 602 is used to provide computing and control capabilities to support the operation of the entire computer device 600.
[0096] The internal memory 604 provides an environment for the operation of the computer program 6032 in the non-volatile storage medium 603. When the computer program 6032 is executed by the processor 602, the processor 602 can be caused to execute the portrait segmentation method.
[0097] The network interface 605 is used for network communication, such as providing the transmission of data information, etc. Those skilled in the art can understand that Figure 6 the structure shown in FIG. is only a block diagram of some structures related to the solution of the present invention, and does not constitute a limitation on the computer device 600 to which the solution of the present invention is applied. The specific computer device 600 may include more or fewer components than those shown in the figure, or combine some components, or have different component arrangements.
[0098] Those skilled in the art can understand that Figure 6 the embodiment of the computer device shown in FIG. does not constitute a limitation on the specific composition of the computer device. In other embodiments, the computer device may include more or fewer components than those shown in the figure, or combine some components, or have different component arrangements. For example, in some embodiments, the computer device may only include a memory and a processor. In such an embodiment, the structures and functions of the memory and the processor are the same as those in Figure 6 the embodiment shown in FIG., and will not be described in detail here.
[0099] It should be understood that in the embodiments of the present invention, the processor 602 may be a central processing unit (CPU), and the processor 602 may also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. Among them, the general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.
[0100] In another embodiment of the present invention, a computer-readable storage medium is provided. The computer-readable storage medium may be a non-volatile computer-readable storage medium. The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the portrait segmentation method of the embodiments of the present invention is implemented.
[0101] The storage medium is a physical, non-transitory storage medium, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a magnetic disk, or an optical disc, etc., which are all physical storage media that can store program codes.
[0102] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the above-described devices, apparatuses, and units can refer to the corresponding processes in the foregoing method embodiments, and will not be elaborated herein.
[0103] As mentioned above, the above are only specific embodiments of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art can easily think of various equivalent modifications or substitutions within the technical scope disclosed by the present invention, and these modifications or substitutions should all be covered by the protection scope of the present invention. Therefore, the protection scope of the present invention should be subject to the protection scope of the claims.
Claims
1. A method for portrait segmentation, characterized in that, Including: Input the image to be segmented into the EfficientNet network for multiple dimensional scaling processes, and respectively output the corresponding dimensionality-reduced images; Input the last two output dimensionality-reduced images into the attention fusion module for feature fusion to obtain a fused image, where the last two output dimensionality-reduced images are 7×7×192 images and 7×7×320 images; Input the fused image into the strip attention module for striping processing to obtain an output image; specifically including: performing three convolution operations on the fused image respectively to obtain a first convolution image, a second convolution image, and a third convolution image; performing a dimensional transpose operation on the first convolution image to obtain a transposed image; performing striping processing and dimensional arrangement processing on the second convolution image, then performing batch processing with the transposed image to obtain an attention matrix; performing striping processing and dimensional transpose processing on the third convolution image, multiplying it with the attention matrix and then performing arrangement processing, and adding it to the fused image to obtain an output image; Segment the target object in the output image and output a target segmentation image.
2. The portrait segmentation method according to claim 1, wherein The step of inputting the last two output dimensionality-reduced images into the attention fusion module for feature fusion to obtain a fused image includes: Performing convolution processing on the 7×7×192 image to obtain a first image; Performing upsampling processing on the 7×7×192 image, then performing convolution, pooling, and activation processing, and performing weight assignment to obtain a second image; Performing a product operation on the first image and the second image, and then performing upsampling processing to obtain a third image; Performing convolution processing on the 7×7×320 image to obtain a fourth image; Performing convolution, pooling, and activation processing on the fourth image, and performing weight assignment to obtain a fifth image; Performing a product operation on the fourth image and the fifth image, and then performing upsampling processing to obtain a sixth image; Performing feature fusion on the third image and the sixth image to obtain a fused image.
3. The portrait segmentation method according to claim 1, wherein: The step of inputting the image to be segmented into the EfficientNet network for multiple dimensional scaling processes and respectively outputting the corresponding dimensionality-reduced images includes: Inputting a 224×224×32 image into the first convolution module for convolution processing to output a 112×112×32 image; Inputting the 112×112×32 image into the second convolution module for convolution processing to output a 56×56×24 image; Inputting the 56×56×24 image into the third convolution module for convolution processing to output a 28×28×40 image; Inputting the 28×28×40 image into the fourth convolution module for convolution processing to output a 28×28×80 image; Inputting the 28×28×80 image into the fifth convolution module for convolution processing to output a 7×7×192 image; Inputting the 7×7×192 image into the sixth convolution module for convolution processing to output a 7×7×320 image.
4. The portrait segmentation method according to claim 3, wherein: The number of convolution layers of the first convolution module is 1, and the convolution kernel is 3×3; The number of convolution layers of the second convolution module is 3, and the convolution kernel is 3×3; The third convolution module has 2 convolution layers and a convolution kernel of 5×5; The fourth convolution module has 3 convolution layers and a convolution kernel of 3×3; The fifth convolution module has 7 convolution layers and a convolution kernel of 5×5; The sixth convolution module has 1 convolution layer and a convolution kernel of 3×3.
5. The portrait segmentation method according to claim 1, wherein Segmenting the target object in the output image and outputting a target segmentation image includes: Obtaining the regional edge information of each object in the output image, and using a semantic segmentation network to segment the target object according to the regional edge information of the target object to obtain a target segmentation image.
6. The portrait segmentation method according to claim 1, wherein It also includes: Optimizing the portrait segmentation network model according to the following loss function: ; Among them, represents a distance function, represents a hyperparameter, represents a region function.
7. A portrait segmentation device, characterized in that, Including: A dimension scaling unit for inputting the image to be segmented into an EfficientNet network for multiple dimension scaling processes and respectively outputting corresponding downscaled images; A feature fusion unit for inputting the last two output downscaled images into an attention fusion module for feature fusion to obtain a fused image, where the last two output downscaled images are 7×7×192 images and 7×7×320 images; A striping processing unit for inputting the fused image into a strip attention module for striping processing to obtain an output image; specifically including: performing three convolution operations on the fused image respectively to obtain a first convolution image, a second convolution image, and a third convolution image; performing a dimension transpose operation on the first convolution image to obtain a transposed image; performing striping processing and dimension arrangement processing on the second convolution image, then performing batch processing with the transposed image to obtain an attention matrix; performing striping processing and dimension transpose processing on the third convolution image, multiplying it with the attention matrix and then performing arrangement processing, and adding it to the fused image to obtain an output image; A segmentation unit for segmenting the target object in the output image and outputting a target segmentation image.
8. A computer device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the portrait segmentation method according to any one of claims 1 to 6.
9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, and when the computer program is executed by the processor, it causes the processor to execute the portrait segmentation method according to any one of claims 1 to 6.
Citation Information
Patent Citations
Shadow detection method based on attention mechanism
CN111639692A
Multi-modal data fusion method based on compound collaborative structure feature recombination network
CN113378989A