Portrait matting processing method and system

By introducing an attention mechanism into the detail branch of the portrait matting model and adjusting the attention weight of the edge pixels of the portrait, the problem of poor matting effect in the existing technology is solved, and higher precision and natural matting effect are achieved.

CN117252888BActive Publication Date: 2026-03-27SO-YOUNG INT INC
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-09-28
Publication Date
2026-03-27

AI Technical Summary

Technical Problem

Existing portrait cutout technologies fail to effectively target the edge pixels of a portrait, resulting in poor cutout effects that cannot meet users' high demands for visual quality.

Method used

By introducing an attention mechanism into the detail branch of the portrait matting model, adjusting the attention weights of the edge pixels of the portrait, and combining semantic and detail information, the matting results are optimized.

Benefits of technology

It improves the accuracy and precision of the matting results, reduces the number of parameters, enhances the model's accuracy, and makes the matting results more natural.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117252888B_ABST
    Figure CN117252888B_ABST
Patent Text Reader

Abstract

The application discloses a portrait matting processing method, which comprises the following steps: obtaining a portrait picture to be processed; processing the portrait picture through a portrait matting model; the portrait matting model extracts semantic information of the portrait picture through a semantic branch, extracts detail information of the portrait picture through a detail branch, and outputs a predicted matting result by fusing the semantic information and the detail information; and in the detail branch, the attention weight of the portrait edge pixels in the portrait picture is adjusted through an attention mechanism. The application also discloses a portrait matting processing system, an electronic device and a computer readable storage medium. Thus, the attention weight of the portrait edge pixels in the model when processing the portrait picture can be adjusted, and the matting effect is improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of image processing, and in particular to a portrait matting processing method and system. BACKGROUND

[0002] In recent years, with the continuous improvement of people's living standards and the continuous increase of the demand for medical insurance of the general public, the medical plastic and cosmetic industry has also developed continuously, and modern medical beauty technology has become increasingly advanced. In user-oriented image processing, face and body part segmentation is a common application scenario. This application scenario requires the use of portrait matting technology.

[0003] Matting refers to the operation of separating a part of a picture or video from the original picture or video to become a separate layer, which can be applied to the fields of portrait blurring, background replacement, and image synthesis. For example, in a 2D portrait stylization product, if the background of the stylized image needs to be replaced, the foreground, i.e. the portrait part, of the stylized image needs to be cut out first, and then fused with a new background. The current image processing software uses neural networks for image processing, which does not require manual operation and can improve the efficiency of image processing and the accuracy of matting.

[0004] However, existing portrait matting technology solutions generally learn the entire image through a network learning model, do not focus on the characteristics of portrait matting, and focus on portrait edge pixels, so the matting effect is not good. SUMMARY

[0005] The main purpose of the present application is to provide a portrait matting processing method and system, an electronic device and a computer readable storage medium, aiming to solve the problem of how to improve the portrait matting effect.

[0006] In a first aspect, an embodiment of the present application provides a portrait matting processing method, which comprises:

[0007] obtaining a portrait picture to be processed;

[0008] processing the portrait picture through a portrait matting model, the portrait matting model extracting semantic information of the portrait picture through a semantic branch, extracting detail information of the portrait picture through a detail branch, and fusing the semantic information and the detail information to output a predicted matting result, wherein in the detail branch, the attention weight of the portrait edge pixel in the portrait picture is adjusted through an attention mechanism.

[0009] In a second aspect, an embodiment of the present application provides a portrait matting processing system, which comprises:

[0010] an acquisition module configured to acquire a portrait picture to be processed;

[0011] The processing module is configured to process the portrait picture by using a portrait matting model. The portrait matting model extracts semantic information of the portrait picture by using a semantic branch, extracts detail information of the portrait picture by using a detail branch, and outputs a predicted matting result by fusing the semantic information and the detail information. In the detail branch, an attention mechanism is used to adjust an attention weight of an edge pixel of a portrait in the portrait picture.

[0012] In a third aspect, an electronic device is provided. The electronic device includes a memory, a processor, and a computer program stored in the memory and executable on the processor. The computer program, when executed by the processor, implements the method described above.

[0013] In a fourth aspect, a computer-readable storage medium is provided. The computer-readable storage medium stores a computer program. The computer program, when executed by a processor, implements the method described above.

[0014] The portrait matting processing method and system, the electronic device, and the computer-readable storage medium provided in the embodiments of the present application introduce an attention mechanism in the detail branch of the portrait matting model, improve the attention weight of the edge pixel of the portrait, reduce the parameter amount on the basis of the original matting effect, improve the model precision, and make the visual effect of the matting result more natural. BRIEF DESCRIPTION OF DRAWINGS

[0015] The accompanying drawings are included to provide a further understanding of the present application, and are incorporated in and constitute a part of this application. It will be appreciated that these drawings depict only some embodiments of the present application and are therefore not to be considered limiting of its scope.

[0016] Figure 1 A flowchart of a portrait matting model training method according to the first embodiment of the present application;

[0017] Figure 2 For Figure 1 A detailed flowchart of step S200;

[0018] Figure 3 A schematic diagram of a portrait picture annotated by using an annotation tool in the present application;

[0019] Figure 4 A network architecture schematic diagram of a network learning model in the present application;

[0020] Figure 5 A flowchart of a portrait matting processing method according to the second embodiment of the present application;

[0021] Figure 6A flowchart of an optional embodiment of the portrait matting processing method based on the second embodiment of the present application;

[0022] Figure 7 A schematic diagram of a white edge in a matting result picture according to the present application;

[0023] Figure 8 A schematic diagram of a white edge in a matting result picture according to the present application; Figure 6 A detailed flowchart of step S304 in the method;

[0024] Figure 9 A schematic diagram of a hardware architecture of an electronic device according to the third embodiment of the present application;

[0025] Figure 10 A schematic diagram of a hardware architecture of an electronic device according to the fifth embodiment of the present application;

[0026] Figure 11 A schematic diagram of a hardware architecture of an electronic device according to the fifth embodiment of the present application;

[0027] Figure 12 A schematic diagram of a hardware architecture of an electronic device according to the fifth embodiment of the present application; DETAILED DESCRIPTION

[0028] In order to make the objects, technical solutions and advantages of the present application clearer, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and should not be used to limit the present application. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative labor fall within the scope of protection of the present application.

[0029] It should be noted that the terms "first", "second", etc. used in the embodiments of the present application are only for the purpose of description and should not be understood as indicating or implying the relative importance of the technical features indicated or the number of the technical features indicated. Therefore, the features with "first" and "second" can explicitly or implicitly include at least one of the features. In addition, the technical solutions of various embodiments can be combined with each other, but it must be based on the fact that a person of ordinary skill in the art can realize it. When the combination of technical solutions appears contradictory or unachievable, it should be considered that the combination of technical solutions does not exist and is not within the scope of protection claimed by the present application.

[0030] Matting refers to the operation of separating a part of a picture or a video from an original picture or an original video to become a separate layer, which can be applied to the fields of portrait blurring, background replacement and image synthesis. The current portrait matting scheme generally learns the entire image through a network learning model, does not focus on the characteristics of portrait matting, and focuses on the edge pixels of the portrait, so the matting effect is not good. Moreover, semantic segmentation mainly provides understanding of each pixel point in the image, and finally outputs pixels with the same semantics as a specific label, without focusing on the actual visual effect. In many application scenarios, users not only pay attention to whether the segmentation result of each semantic region is correct, but also put forward higher requirements on whether the observed visual effect is more friendly.

[0031] Therefore, the embodiment of the present application provides a new portrait matting processing scheme. The implementation idea of the new portrait matting processing scheme provided by the embodiment of the present application includes: obtaining a portrait picture to be processed; extracting semantic information of the portrait picture through a semantic branch, extracting detail information of the portrait picture through a detail branch, and fusing the semantic information and the detail information to output a predicted matting result. In the detail branch, the attention weight of the edge pixel of the portrait in the portrait picture is adjusted through an attention mechanism, so that the accuracy of the matting result and the fineness of the matting result are improved under the condition of reducing the parameter amount.

[0032] Under the guidance of the idea of the present application, the technical scheme of the present application will be specifically described in combination with various embodiments.

[0033] Embodiment one

[0034] As shown in Figure 1 , a flowchart of a portrait matting model training method proposed by the first embodiment of the present application. It can be understood that the flowchart in the embodiment of the present method is not used to limit the order of executing steps. According to the needs, some steps in the flowchart can also be added or deleted. The method will be described below with the server as the execution subject.

[0035] The method includes the following steps:

[0036] S200, a plurality of portrait picture samples are collected to generate a training data set.

[0037] In the embodiment, the training data set used for training the portrait matting model is a plurality of portrait picture samples. The plurality of portrait picture samples are obtained by preprocessing a plurality of portrait pictures collected in advance, and each original portrait picture can be processed into a group of different portrait picture samples.

[0038] Specifically, further referring to Figure 2, a detailed flowchart of step S200. It can be understood that the flowchart does not limit the order of execution of the steps. According to the needs, some steps in the flowchart can be added or deleted. In this embodiment, the step S200 specifically includes:

[0039] S2000, collecting multiple portrait pictures.

[0040] The portrait picture can be a photo collected by shooting a portrait, or a portrait picture obtained from a pre-stored image database.

[0041] S2002, respectively labeling the multiple portrait pictures by a labeling tool, and cutting multiple matting samples corresponding to the multiple portrait pictures according to the labeling.

[0042] The labeling tool is used to label the portrait edge of the portrait picture, so as to more accurately mat the portrait picture according to the labeling. For example, the labeling tool can be labelme. The labelme is an online Javascript image labeling tool, which can be used to create customized labeling tasks or perform image labeling. The labelme can label images in the forms of polygon, rectangle, circle, polyline, line segment and point, which are used for target detection, image segmentation and other tasks, or label images in the form of flag, which is used for image classification and cleaning tasks. In this embodiment, first, the target (portrait) is outlined by the polygon tool of the labelme, and then the label is created after the outline is completed, and the frame of the polygon is fine-tuned, so as to cut out the portrait part. As shown in Figure 3 , it is a schematic diagram of labeling a portrait picture by a labeling tool. In Figure 3 , the dots and line segments at the edge of the portrait are the labeling. In other embodiments, other tools or methods can be used to label the portrait picture, which will not be described here.

[0043] S2004, cropping and data enhancement processing are performed on each portrait picture to obtain a group of different portrait picture samples.

[0044] In this embodiment, in order to enrich the samples, various forms of cropping and different data enhancement processing can be performed on each portrait picture, that is, the diversity of the portrait picture is changed, so that a group of different portrait picture samples are obtained from one portrait picture, and the group of portrait picture samples correspond to the same portrait matting result. Therefore, the training data set can be expanded, and the portrait matting model can be trained from multiple aspects to improve the model performance.

[0045] Back to Figure 1S202, training a network learning model according to the training data set.

[0046] The plurality of groups of portrait picture samples in the training data set are input into a pre-trained model, and after repeated training, a trained portrait matting model is obtained. The pre-trained model can be a general network learning model. The pre-trained model is usually a model trained on a large data set, such as the BERT model in natural language processing (NLP), the ResNet, SwinTransformer model, and PP-Matting model in computer vision (CV), etc. Through the pre-trained model, the "knowledge" learned on one task can be applied to a new task or a new data set. Using the pre-trained model can provide better performance than building a model from scratch.

[0047] In this embodiment, the network learning model extracts semantic information of each portrait picture sample in the training data set through a semantic branch, extracts detail information of each portrait picture sample in the training data set through a detail branch, and fuses the semantic information and the detail information to obtain a predicted matting result. It is worth noting that in order to improve the attention of the model to the edge details, an attention mechanism is also introduced in the detail branch, which adjusts the attention weight of the portrait edge pixels in the portrait picture sample through the attention mechanism. The adjustment specifically refers to increasing the attention weight of the portrait edge pixels in the portrait picture sample and reducing the attention weight of other pixels.

[0048] The introduction of the attention mechanism specifically refers to the introduction of an attention module in the detail branch. The attention module first divides the input tensor (Input Tensor) into multiple groups from the channel through a first submodule, each group performs convolution with different kernel sizes to extract information of different scales, and the kernel size of each group increases in turn. Then the output of the first submodule is extracted through a second submodule to obtain the channel attention value of each group, the purpose of which is to obtain the attention weight of the feature maps of different scales. Through such a method, the attention module fuses the context information of different scales and produces better pixel-level attention. Finally, the channel attention value of each group is spliced and normalized, and the output of the first submodule is weighted. The purpose of the attention module is to increase the spatial information of different scales to enrich the feature space, while considering the dependency relationship between local regions and distant regions, and reducing the computational complexity.

[0049] Specifically, as shown in Figure 4 , it is a network architecture diagram of the network learning model. In Figure 4In the semantic branch, the portrait picture sample input into the model is processed by multi-layer down-sampling, pyramid pooling module and multi-layer up-sampling to extract high-level semantic information of the picture. In an optional embodiment, the multi-layer down-sampling changes the number of channels to 1 / 2, 1 / 4, 1 / 8, 1 / 16 and 1 / 32 of the original number of channels respectively; the pyramid pooling module provides strong extraction function of global context for semantic segmentation; and the multi-layer up-sampling changes the number of channels to 1 / 16, 1 / 8, 1 / 4, 1 / 2 and 1 / 1 of the original number of channels respectively to obtain the semantic map of the portrait picture sample.

[0050] Further, the edge transition area of the portrait picture sample is extracted by the high-resolution detail branch designed by guiding flow to obtain a detail map. Meanwhile, the attention module is used to improve the recognition accuracy and efficiency of detail sampling. Specifically, the 1 / 2 channel (first layer down-sampling result) and 1 / 32 channel (last layer down-sampling result) in the down-sampling process are spliced, and then the first processing is performed by the attention module; the output result of the first processing is multiplied by the 1 / 16 channel (first layer up-sampling result) in the up-sampling process, and then the second processing is performed by the attention module; the output result of the second processing is multiplied by the 1 / 4 channel (intermediate layer up-sampling result) in the up-sampling process, and then the third processing is performed by the attention module; the output result of the third processing is multiplied by the 1 / 1 channel (last layer up-sampling result) in the up-sampling process, and then the fourth processing is performed by the attention module, and finally the detail map is obtained. Before the splicing or matrix multiplication operation, if the levels (i.e. the number of channels) of the two sides of the operation do not match, the up-sampling processing is needed to make the levels of the two sides of the operation match before the splicing or matrix multiplication operation. The level matching refers to that the number of channels of the two sides of the operation is the same.

[0051] Finally, the fusion of the semantic map and the detail map is realized by the fusion module to obtain the final predicted matting result. Figure 4 In the semantic branch, the portrait picture sample input into the model is processed by multi-layer down-sampling, pyramid pooling module and multi-layer up-sampling to extract high-level semantic information of the picture. In an optional embodiment, the multi-layer down-sampling changes the number of channels to 1 / 2, 1 / 4, 1 / 8, 1 / 16 and 1 / 32 of the original number of channels respectively; the pyramid pooling module provides strong extraction function of global context for semantic segmentation; and the multi-layer up-sampling changes the number of channels to 1 / 16, 1 / 8, 1 / 4, 1 / 2 and 1 / 1 of the original number of channels respectively to obtain the semantic map of the portrait picture sample.

[0052] Back to Figure 1S204, whether the network learning model reaches the preset training target according to the loss function. If the network learning model reaches the training target, the process ends, and a trained portrait matting model is obtained. If the network learning model does not reach the training target, step S206 is performed.

[0053] The loss function is used to measure the deviation between the predicted value made by the model and the true value. In this embodiment, the loss function is used to measure the error between the predicted matting result output by the network learning model and the true portrait matting result corresponding to the portrait picture sample.

[0054] In an optional embodiment, the training process of the network learning model includes three parts of loss function. The first part is the cross-entropy loss function of the semantic branch, referred to as the first loss function. The second part is the loss function of the detail branch, referred to as the second loss function. The third part is the fusion loss function, referred to as the third loss function. The calculation formula of the final loss function of the model is:

[0055] L = λ1L s + λ2L d + λ3L f

[0056] Wherein, L s represents the first loss function, L d represents the second loss function, L f represents the third loss function, and λ1, λ2, λ3 are the weights of the above three loss functions, respectively.

[0057] When the final loss function is lower than the preset threshold, it indicates that the network learning model has reached the training target, and the trained network learning model can be used as the final portrait matting model. When the final loss function is not lower than the preset threshold, it indicates that the network learning model has not reached the training target, and further training needs to be continued.

[0058] S206, adjust the model parameters and return to the above step S202 to repeat the model training until the training target is reached.

[0059] In this embodiment, when the network learning model does not reach the training target, the optimizer and the learning rate scheduler can be used to adjust the corresponding model parameters. The learning rate is a tuning parameter in the optimization algorithm, which can determine the step size in each iteration, so that the loss function converges to the minimum value.

[0060] The optimizer can be an AdamW optimizer, which combines the Adam optimizer (Adaptive Moment Estimation Optimizer) and the concept of weight decay. The Adam optimizer is a gradient-based optimization algorithm that can adaptively adjust the learning rate. The learning rate is updated based on the historical gradient information and the average value of each parameter, so that a larger learning rate is used at the beginning of training to quickly converge, and a smaller learning rate is used at the end of training to more accurately find the minimum value of the loss function. The term of weight decay is added to the Adam optimizer to regularize the parameters during updating, which can control the complexity of the model and prevent the neural network from overfitting the training data.

[0061] The learning rate scheduler adopts Cosine Annealing Warm Restarts. The cosine annealing algorithm can reduce the learning rate through a cosine function. In the cosine function, the cosine value first slowly decreases, then accelerates, and then slowly decreases as x increases. This descending pattern can cooperate with the learning rate to make the model jump out of the local optimal solution, thereby training a better model. The learning rate formula of the learning rate scheduler is as follows:

[0062]

[0063] wherein η min is the minimum learning rate, η max is the initial learning rate, T cur is the number of training epochs after the last learning rate reset, and T i represents the number of training epochs after which the learning rate is reset. When T cur = T i , set η t = η min ; when the learning rate is reset, T cur = 0, set η t = η max .

[0064] The optimizer can avoid using a separate learning rate scheduler, but instead choose to directly embed the learning rate optimization into the optimizer itself. After adjusting the parameters of the network learning model by the optimizer and the learning rate scheduler, return to the above step S202, repeat the model training according to the training data set and the adjusted model, until the network learning model reaches the training target, that is, the final loss function is lower than the preset threshold, and the trained portrait matting model is obtained.

[0065] The portrait matting model training method proposed in this embodiment improves the attention weight of the portrait edge pixels in the portrait picture sample by introducing an attention mechanism in the detail branch of the network learning model, reduces the parameter quantity on the basis of the original matting effect, and improves the model precision. In addition, when the model training does not reach the target, by adjusting the learning rate of each parameter, the model can learn to pay more attention to the key pixels on the edge, and a better model can be trained. Moreover, compared with the classic semantic segmentation task, the portrait matting model trained in this embodiment can make the visual effect of the matting result more natural.

[0066] Embodiment Two

[0067] As Figure 5 shown, a flowchart of a portrait matting processing method according to the second embodiment of the present application is shown. It can be understood that the flowchart in the method embodiment is not used to limit the order of executing steps. According to the needs, some steps in the flowchart can be added or deleted. The portrait matting processing method can be executed by the client or the server, which is not limited here.

[0068] The method comprises the following steps:

[0069] S300, obtaining a portrait picture to be processed.

[0070] The portrait picture to be processed can be a portrait picture currently taken by a user, a portrait picture obtained from a network, a database, etc., which needs to be processed by portrait matting.

[0071] S302, processing the portrait picture by a portrait matting model to output a predicted matting result.

[0072] The portrait picture to be processed is input into the trained portrait matting model, and the portrait picture is automatically processed by the portrait matting model to output a corresponding predicted matting result. The portrait matting model is obtained by training a network learning model according to a plurality of portrait picture samples, and the learning rate of each parameter of the network learning model is adjusted by an optimizer and a learning rate scheduler during the training process. The specific training process of the portrait matting model can be referred to the description in the first embodiment, which will not be repeated here.

[0073] In the embodiment, the portrait matting model extracts semantic information of the portrait picture through a semantic branch, extracts detail information of the portrait picture through a detail branch, and obtains the predicted matting result by fusing the semantic information and the detail information. The attention mechanism is introduced in the detail branch, and the attention weight of the portrait edge pixel in the portrait picture is adjusted through the attention mechanism. The adjustment specifically refers to increasing the attention weight of the portrait edge pixel in the portrait picture and reducing the attention weight of other pixels.

[0074] The attention mechanism is introduced specifically by introducing an attention module in the detail branch. The attention module first divides the input tensor into multiple groups from the channel by a first submodule, each group performs convolution with different kernel sizes to extract information of different scales, and the convolution kernel size of each group increases in turn. Then the output of the first submodule is extracted through a second submodule to obtain the channel attention value of each group, the purpose of which is to obtain the attention weight of the feature map of different scales. Through such a method, the attention module fuses the context information of different scales and produces better pixel-level attention. Finally, the channel attention value of each group is spliced and normalized, and the output of the first submodule is weighted. The purpose of the attention module is to increase the spatial information of different scales to enrich the feature space, while considering the dependence between local regions and distant regions, and reducing the computational complexity.

[0075] The portrait picture input into the portrait matting model is processed by multi-layer down-sampling, pyramid pooling module and multi-layer up-sampling in the semantic branch to extract high-level semantic information of the picture. In an optional embodiment, the multi-layer down-sampling changes the channel number to 1 / 2, 1 / 4, 1 / 8, 1 / 16 and 1 / 32 of the original channel number respectively; the pyramid pooling module provides strong extraction function of global context for semantic segmentation; and the multi-layer up-sampling changes the channel number to 1 / 16, 1 / 8, 1 / 4, 1 / 2 and 1 / 1 of the original channel number respectively to obtain the semantic map of the portrait picture sample.

[0076] Further, the edge transition area of the portrait picture sample is extracted by guiding the flow design to gradually guide the high-resolution detail branch to obtain a detail picture. Meanwhile, the attention module is used to improve the recognition accuracy and efficiency of the detail sampling. Specifically, the 1 / 2 channel (first layer down-sampling result) and the 1 / 32 channel (last layer down-sampling result) in the down-sampling process are spliced, and then the first processing is performed by the attention module; the output result of the first processing is multiplied by the 1 / 16 channel (first layer up-sampling result) in the up-sampling process, and then the second processing is performed by the attention module; the output result of the second processing is multiplied by the 1 / 4 channel (intermediate layer up-sampling result) in the up-sampling process, and then the third processing is performed by the attention module; the output result of the third processing is multiplied by the 1 / 1 channel (last layer up-sampling result) in the up-sampling process, and then the fourth processing is performed by the attention module, and finally the detail picture is obtained. Before the splicing or matrix multiplication operation, if the levels (i.e., the number of channels) of the two sides of the operation do not match, the up-sampling processing is needed to be performed first to match the levels of the two sides of the operation, and then the splicing or matrix multiplication operation is performed. The level matching refers to that the number of channels of the two sides of the operation is the same.

[0077] Finally, the fusion of the semantic picture and the detail picture is realized by the fusion module to obtain the final predicted cutout result of the portrait picture.

[0078] In an optional embodiment, as shown in Figure 6 the portrait cutout processing method further comprises:

[0079] S304, the predicted cutout result output by the portrait cutout model is edge-removed to obtain a final cutout picture.

[0080] Generally, the predicted cutout result output by the portrait cutout model may have a white edge or a black edge around the portrait area. As shown in Figure 7 , it is a white edge diagram in a cutout result picture. Therefore, in this embodiment, in order to make the portrait cutout more clear and realistic, the predicted cutout result needs to be further edge-removed to remove the white edge or the black edge.

[0081] Specifically, further referring to Figure 8 , it is a detailed flowchart of the step S304. It can be understood that the flowchart is not used to limit the order of the execution steps. According to the needs, some steps in the flowchart can be added or deleted. In this embodiment, the step S304 specifically comprises:

[0082] S3040, the RGB (red, green and blue channels) values of all pixel points in the predicted cutout result picture are obtained.

[0083] S3042, detect the white or black pixels in the predicted cutout result image based on the RGB values.

[0084] Since white or black borders usually appear in the predicted image cutout results, it is first necessary to detect white or black pixels based on RGB values.

[0085] S3044, determine the white or black border area based on the white or black pixels and the position of the portrait edge in the predicted cutout result image.

[0086] Specifically, white pixels located at the edge of the portrait in the predicted matting result image belong to the white edge area of ​​the predicted matting result image. Alternatively, black pixels located at the edge of the portrait in the predicted matting result image belong to the black edge area of ​​the predicted matting result image.

[0087] S3046, Remove the white or black border areas to obtain the final cutout image.

[0088] The portrait matting method proposed in this embodiment can improve the model output by introducing an attention mechanism into the detail branch of the portrait matting model, thereby increasing the attention weight of edge pixels in the portrait image to be processed. Furthermore, further edge removal processing of the predicted matting results output by the model can make the portrait matting clearer and more realistic.

[0089] Example 3

[0090] like Figure 9 The diagram shown illustrates the hardware architecture of an electronic device 20 according to a third embodiment of this application. In this embodiment, the electronic device 20 may include, but is not limited to, a memory 21, a processor 22, and a network interface 23, which are interconnected via a system bus. It should be noted that... Figure 9 Only the electronic device 20 with components 21-23 is shown; however, it should be understood that implementation of all shown components is not required, and more or fewer components may be implemented alternatively. In this embodiment, the electronic device 20 may be a server.

[0091] The memory 21 includes at least one type of readable storage medium, including flash memory, hard disk, multimedia card, card-type memory (e.g., SD or DX memory), random access memory (RAM), static random access memory (SRAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), programmable read-only memory (PROM), magnetic memory, magnetic disk, optical disk, etc. In some embodiments, the memory 21 may be an internal storage unit of the electronic device 20, such as the hard disk or memory of the electronic device 20. In other embodiments, the memory 21 may also be an external storage device of the electronic device 20, such as a plug-in hard disk, smart media card (SMC), secure digital (SD) card, flash card, etc., equipped on the electronic device 20. Of course, the memory 21 may include both the internal storage unit and the external storage device of the electronic device 20. In this embodiment, the memory 21 is typically used to store the operating system and various application software installed on the electronic device 20, such as the program code of the portrait matting model training system 60. In addition, the memory 21 can also be used to temporarily store various types of data that have been output or will be output.

[0092] In some embodiments, the processor 22 may be a central processing unit (CPU), controller, microcontroller, microprocessor, or other data processing chip. The processor 22 is typically used to control the overall operation of the electronic device 20. In this embodiment, the processor 22 is used to run program code stored in the memory 21 or process data, such as running the portrait matting model training system 60.

[0093] The network interface 23 may include a wireless network interface or a wired network interface, which is typically used to establish a communication connection between the electronic device 20 and other electronic devices.

[0094] Example 4

[0095] like Figure 10 The diagram shown is a modular schematic of a portrait matting model training system 60 according to the fourth embodiment of this application. The portrait matting model training system 60 can be divided into one or more program modules. One or more program modules are stored in a storage medium and executed by one or more processors to complete the embodiments of this application. The program module referred to in the embodiments of this application refers to a series of computer program instruction segments capable of performing specific functions. The following description will specifically introduce the functions of each program module in this embodiment.

[0096] In the embodiment, the portrait matting model training system 60 comprises:

[0097] The acquisition module 600 is configured to acquire a plurality of portrait picture samples to generate a training data set.

[0098] Specifically, a plurality of portrait pictures are first acquired, and then the plurality of portrait pictures are labeled respectively by using a labeling tool. A plurality of matting samples corresponding to the plurality of portrait pictures are obtained by cutting according to the labels. Finally, each portrait picture is cropped and subjected to data enhancement processing to obtain a plurality of different portrait picture samples.

[0099] The training module 602 is configured to train a network learning model according to the training data set.

[0100] The plurality of portrait picture samples in the training data set are input into a pre-trained model, and a trained portrait matting model can be obtained through repeated training. The pre-trained model can be a general network learning model.

[0101] In the embodiment, the network learning model extracts semantic information of each portrait picture sample in the training data set through a semantic branch, extracts detail information of each portrait picture sample in the training data set through a detail branch, and obtains a predicted matting result by fusing the semantic information and the detail information. It is worth noting that, in order to improve the attention of the model to the edge details, an attention mechanism is introduced in the detail branch, and the attention weight of the portrait edge pixels in the portrait picture sample is adjusted through the attention mechanism. The adjustment specifically refers to increasing the attention weight of the portrait edge pixels in the portrait picture sample and reducing the attention weight of other pixels.

[0102] The introduction of the attention mechanism specifically refers to the introduction of an attention module in the detail branch. The attention module first divides the input tensor into a plurality of groups from the channel by using a first submodule, each group is subjected to convolution with different kernel sizes to extract information of different scales, and the kernel size of each group is increased in turn. The output of the first submodule is extracted through a second submodule to obtain the channel attention value of each group, and the purpose is to obtain the attention weight of the feature maps of different scales. Through such a method, the attention module fuses the context information of different scales and produces better pixel-level attention. Finally, the channel attention weight of each group is spliced and normalized, and the output of the first submodule is weighted. The purpose of the attention module is to increase the spatial information of different scales to enrich the feature space, consider the dependence relationship between local regions and distant regions, and reduce the calculation amount.

[0103] The test module 604 is configured to test whether the network learning model reaches a preset training target according to the loss function. If the network learning model reaches the training target, a trained portrait matting model is obtained.

[0104] In this embodiment, the loss function is used to measure the error between the predicted matting result output by the network learning model and the real portrait matting result corresponding to the portrait picture sample. When the final loss function is lower than a preset threshold, it indicates that the network learning model has reached the training target, and the trained network learning model can be used as a final portrait matting model.

[0105] The optimization module 606 is configured to adjust the model parameters when the network learning model does not reach the training target. Then the model training is repeated by the training module 602 until the training target is reached.

[0106] In this embodiment, when the network learning model does not reach the training target, an optimizer and a learning rate scheduler can be used to adjust the corresponding model parameters. The optimizer can be an AdamW optimizer, which combines the concepts of Adam optimizer and weight decay. The AdamW optimizer can adaptively adjust the learning rate, update the learning rate according to the historical gradient information and average value of each parameter, control the complexity of the model, and prevent the neural network from overfitting the training data. The learning rate scheduler uses cosine annealing warm restart. The cosine annealing algorithm can reduce the learning rate through a cosine function. This descending mode can cooperate with the learning rate to make the model jump out of the local optimal solution, so as to train a better model.

[0107] The specific implementation process of each module function can be referred to the description in the first embodiment above, which will not be repeated here.

[0108] The portrait matting model training system proposed in this embodiment can improve the attention weight of the edge pixels of the portrait in the portrait picture sample by introducing an attention mechanism in the detail branch of the network learning model, reduce the parameter amount on the basis of the original matting effect, and improve the model precision. In addition, when the model training does not reach the target, adjusting the learning rate of each parameter can make the model pay more attention to the key pixels on the edge, and train a better model. Moreover, compared with the classic semantic segmentation task, the portrait matting model trained in this embodiment can make the visual effect of the matting result more natural.

[0109] Embodiment five

[0110] As Figure 11As shown, a hardware architecture diagram of an electronic device 30 is provided in the fifth embodiment of the present application. In this embodiment, the electronic device 30 can include, but is not limited to, a memory 31, a processor 32, and a network interface 33, which can be communicatively connected to each other through a system bus. It should be noted that, Figure 11 Only the electronic device 30 with components 31-33 is shown, but it should be understood that all of the illustrated components are not required, and that more or less components can be implemented instead. In this embodiment, the electronic device 30 can be a server or a client.

[0111] The memory 31 includes at least one type of readable storage medium, including a flash memory, a hard disk, a multimedia card, a card-type memory (e.g., an SD or DX memory, etc.), a random access memory (RAM), a static random access memory (SRAM), a read-only memory (ROM), an electrically erasable programmable read-only memory (EEPROM), a programmable read-only memory (PROM), a magnetic memory, a magnetic disk, an optical disk, etc. In some embodiments, the memory 31 can be an internal storage unit of the electronic device 30, such as a hard disk or a memory of the electronic device 30. In other embodiments, the memory 31 can also be an external storage device of the electronic device 30, such as a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card, etc. Of course, the memory 31 can include both an internal storage unit and an external storage device of the electronic device 30. In this embodiment, the memory 31 is generally used to store an operating system and various application software installed in the electronic device 30, such as program codes of the portrait matting processing system 70, etc. In addition, the memory 31 can also be used to temporarily store various data that have been output or will be output.

[0112] The processor 32 can be a CPU, a controller, a microcontroller, a microprocessor, or other data processing chip in some embodiments. The processor 32 is generally used to control the overall operation of the electronic device 30. In this embodiment, the processor 32 is used to run program codes or process data stored in the memory 31, such as running the portrait matting processing system 70, etc.

[0113] The network interface 33 can include a wireless network interface or a wired network interface, and is generally used to establish a communication connection between the electronic device 30 and other electronic devices.

[0114] Embodiment Six

[0115] As Figure 12As shown, a module schematic diagram of a portrait matting processing system 70 is provided in the sixth embodiment of the present application. The portrait matting processing system 70 can be divided into one or more program modules, and the one or more program modules are stored in a storage medium and executed by one or more processors to complete the embodiments of the present application. The program module referred to in the embodiments of the present application refers to a series of computer program instruction segments capable of completing a specific function, and the functions of the program modules of the embodiments will be described in detail below.

[0116] In the embodiment, the portrait matting processing system 70 includes:

[0117] An acquisition module 700 is configured to acquire a portrait picture to be processed.

[0118] A processing module 702 is configured to process the portrait picture by using a portrait matting model, and output a predicted matting result.

[0119] In the embodiment, the portrait matting model extracts semantic information of the portrait picture by using a semantic branch, extracts detail information of the portrait picture by using a detail branch, and fuses the semantic information and the detail information to obtain the predicted matting result. In the detail branch, an attention module is introduced, which can adjust an attention weight of a portrait edge pixel in the portrait picture. The adjustment specifically refers to increasing the attention weight of the portrait edge pixel in the portrait picture and decreasing an attention weight of other pixels.

[0120] The specific implementation process of each module function can be referred to the description in the first embodiment above, and will not be described here.

[0121] The portrait matting processing system provided in the embodiment can improve the attention weight of a portrait edge pixel in a portrait picture to be processed by introducing an attention mechanism in a detail branch of the portrait matting model, and improve the output effect of the model.

[0122] Embodiment Seven

[0123] The present application also provides another implementation, that is, a computer readable storage medium storing a portrait matting model training program or a portrait matting processing program, and the portrait matting model training program or the portrait matting processing program can be executed by at least one processor to enable the at least one processor to perform the steps of the portrait matting model training method or the portrait matting processing method as described above.

[0124] In this embodiment, the computer readable storage medium includes a flash memory, a hard disk, a multimedia card, a card-type memory (e.g., an SD or DX memory, etc.), a random access memory (RAM), a static random access memory (SRAM), a read-only memory (ROM), an electrically erasable programmable read-only memory (EEPROM), a programmable read-only memory (PROM), a magnetic memory, a magnetic disk, an optical disk, etc. In some embodiments, the computer readable storage medium can be an internal storage unit of the computer device, such as a hard disk or a memory of the computer device. In other embodiments, the computer readable storage medium can also be an external storage device of the computer device, such as a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card, etc. Of course, the computer readable storage medium can also include both the internal storage unit and the external storage device of the computer device. In this embodiment, the computer readable storage medium is usually used to store an operating system and various application software installed on the computer device, such as program codes of the portrait matting model training method or the portrait matting processing method in the embodiments. In addition, the computer readable storage medium can also be used to temporarily store various data that have been output or will be output.

[0125] It should be noted that, in this document, the terms "comprising", "containing" or any other variant thereof are intended to cover non-exclusive inclusions, so that a process, method, article or device that includes a series of elements not only includes those elements, but also includes other elements not explicitly listed or inherent to such a process, method, article or device. Without more limitations, the element defined by the statement "comprising a" does not exclude the presence of another identical element in the process, method, article or device that includes the element.

[0126] The serial numbers of the embodiments of the present application described above are only for description, and do not represent the advantages or disadvantages of the embodiments.

[0127] Obviously, those skilled in the art should understand that each module or step of the above-mentioned embodiments of the present application can be realized by a general computing device, which can be concentrated on a single computing device or distributed on a network composed of multiple computing devices, and can be realized by program codes executable by a computing device, so that they can be stored in a storage device and executed by a computing device, and in some cases, the steps shown or described can be executed in different order, or they can be manufactured into individual integrated circuit modules, or multiple modules or steps among them can be manufactured into a single integrated circuit module. Thus, the embodiments of the present application are not limited to any specific combination of hardware and software.

[0128] The above merely preferred embodiments of the present application are not intended to limit the patent scope of the present application, and any equivalent structure or equivalent process transformation, or direct or indirect application in other related technical fields, which are made by using the content of the present application specification and drawings, are also included in the patent protection scope of the present application.

Claims

1. A method for processing portrait cutout, characterized in that, The method includes: Obtain the portrait image to be processed; The portrait image is processed by a portrait matting model. The portrait matting model extracts semantic information of the portrait image through a semantic branch and extracts detailed information of the portrait image through a detail branch. The semantic information and the detailed information are then fused to output a predicted matting result. In the detail branch, the attention weights of the edge pixels of the portrait image are adjusted through an attention mechanism. The step of extracting detailed information from the portrait image through detail branches includes: An attention module is introduced into the detail branch. The first and last downsampling results of the semantic branch are concatenated and then processed for the first time by the attention module. The output of the first processing is multiplied by the first layer upsampling result of the semantic branch, and then processed a second time through the attention module. The output of the second processing is multiplied by the intermediate layer upsampling result of the semantic branch, and then processed a third time through the attention module. After performing matrix multiplication between the output of the third processing and the last layer upsampling result of the semantic branch, the fourth processing is performed through the attention module to finally obtain the detail image of the portrait image.

2. The portrait cutout processing method according to claim 1, characterized in that, The step of adjusting the attention weights of the edge pixels of the portrait image through an attention mechanism includes: The attention module divides the input tensor into multiple groups from the channels through the first submodule, and performs convolution on each group with different kernel sizes to extract information at different scales. The output of the first submodule is used to extract the channel attention value of each group through the second submodule. The channel attention values ​​of each group are concatenated and normalized, and the output of the first submodule is weighted.

3. The portrait cutout processing method according to claim 1, characterized in that, The introduction of an attention module in the detail branch also includes: Before the concatenation or matrix multiplication operation, if the levels of the two operations do not match, an upsampling process is performed to match the levels before the operation.

4. The portrait cutout processing method according to claim 1, characterized in that, The method further includes: The predicted matting result output by the portrait matting model is processed to remove edges, resulting in the final matted image.

5. The portrait cutout processing method according to claim 4, characterized in that, The process of removing edges from the predicted matting result output by the portrait matting model includes: Obtain the red, green, and blue channel values ​​of all pixels in the cutout image; Based on the red, green and blue channel values, detect the white or black pixels in the cutout result image; The white or black border area is determined based on the white or black pixels and the position of the human figure edge in the cutout result image; Remove the white or black border areas to obtain the final cutout image.

6. The portrait cutout processing method according to claim 1, characterized in that, The portrait matting model is obtained by training a network learning model based on multiple sets of portrait image samples. During the training process, the learning rate of each parameter of the network learning model is adjusted by an optimizer and a learning rate scheduler.

7. A portrait cutout processing system, characterized in that, The system includes: The acquisition module is used to acquire portrait images to be processed. The processing module is used to process the portrait image through a portrait matting model. The portrait matting model extracts the semantic information of the portrait image through a semantic branch and extracts the detail information of the portrait image through a detail branch. It then fuses the semantic information and the detail information to output a predicted matting result. In the detail branch, the attention weights of the edge pixels of the portrait image are adjusted through an attention mechanism. The processing module is further configured to introduce an attention module into the detail branch, concatenate the first-level downsampling result and the last-level downsampling result of the semantic branch, and then perform a first processing through the attention module. The output of the first processing is then multiplied by the matrix of the first-level upsampling result of the semantic branch, and then performed a second processing through the attention module. The output of the second processing is then multiplied by the matrix of the intermediate-level upsampling result of the semantic branch, and then performed a third processing through the attention module. The output of the third processing is then multiplied by the matrix of the last-level upsampling result of the semantic branch, and then performed a fourth processing through the attention module, finally obtaining the detail image of the portrait image.

8. An electronic device, characterized in that, The electronic device includes: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the computer program, when executed by the processor, implements the method as described in any one of claims 1 to 6.

9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed by a processor, implements the method as described in any one of claims 1 to 6.

10. A computer program product, comprising a computer program, characterized in that, When executed by a processor, the computer program implements the steps of the method according to any one of claims 1 to 6.

Citation Information

Patent Citations

  • Face segmentation method and device and computer readable storage medium

    CN112330696A