Method for determining matting parameters, matting method, device and electronic equipment
Patent Information
- Application Number
- CN202611115773.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-24
- Publication Date
- 2026-09-01
AI Technical Summary
[0004]这种方式不仅操作繁琐、效率低下,而且由于用户的视觉经验存在差异,很难稳定得到较佳的抠图参数,使得用户的抠图质量难以得到稳定保障,对于不具备专业图像处理经验的用户来说,更难快速调出合适的参数,体验较差
本申请实施例提供的抠图参数的确定方法,获取了待抠图图像后,基于预先训练的分割模型确定所述待抠图图像中各个像素对应的真值分割信息,由于所述真值分割信息用于指示所对应的像素为背景或前景,所以,真值分割信息能够比较准确的反映待抠图图像的各个像素为前景还是背景,再确定待优化的抠图参数,即确定初始待优化的抠图参数,该抠图参数用于控制图像整体的抠图强度,接着按所确定的抠图参数对所述待抠图图像进行抠图渲染,得到效果分割信息,之后以减小所述效果分割信息与真值分割信息之间的损失为优化目标,调整抠图参数,最终得到适配当前待抠图图像的目标抠图参数。整个过程无需用户手动调试参数,能够自动完成抠图参数的优化确定,既提升了抠图参数确定的效率,也能保障参数确定的准确性,让不同经验水平的用户都能稳定获得合适的抠图参数,保障最终的抠图效果。
Smart Images

Figure CN122672703A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of image processing technology, specifically to a method for determining matting parameters, a matting method, an apparatus, an electronic device, and a computer-readable storage medium. Background Technology
[0002] With the rapid development of image and video technologies, the demand for image processing is increasing daily. Background removal, as an important operation, is widely used in many scenarios such as image editing, content creation, and green screen cutout for live streaming. The quality of background removal is highly dependent on the parameters used. These parameters control the leniency of background color discrimination and determine the overall strength of the background removal. If the parameters are not set properly, problems such as the foreground being misjudged as background or excessive background residue can easily occur, affecting the final quality of the background removal.
[0003] Traditional methods of determining image cutout parameters mostly rely on manual adjustments by the user. Users need to repeatedly adjust the parameter values based on the preview image until they are satisfied with the visual effect.
[0004] This method is not only cumbersome and inefficient, but also makes it difficult to consistently obtain optimal cutout parameters due to differences in users' visual experience. This results in inconsistent cutout quality, and for users without professional image processing experience, it is even more difficult to quickly find suitable parameters, leading to a poor user experience. Therefore, how to automatically determine appropriate cutout parameters and improve the efficiency and accuracy of parameter determination has become an urgent problem to be solved. Summary of the Invention
[0005] This application provides a method, apparatus, electronic device, and computer-readable storage medium for determining image matting parameters. It can automatically determine suitable matting parameters, effectively improving the efficiency and accuracy of parameter determination. It eliminates the need for repeated manual adjustments by the user, lowering the barrier to entry for parameter adjustment. Simultaneously, it can more stably output matting parameters adapted to the current image to be matted, resulting in more stable matting effects and improved matting quality and user experience. The specific solution is as follows: Firstly, this application provides a method for determining image matting parameters, the method comprising: Obtain the image to be cut out; Based on a pre-trained segmentation model, ground truth segmentation information is determined for each pixel in the image to be cut out. The ground truth segmentation information is used to indicate whether the corresponding pixel is the background or the foreground. Determine the matting parameters to be optimized, which are used to control the overall matting intensity of the image; The image to be cut out is rendered according to the cutout parameters to obtain the effect segmentation information corresponding to each pixel in the image to be cut out; With the optimization objective of reducing the loss between the effect segmentation information and the ground truth segmentation information, the matting parameters are adjusted to obtain the target matting parameters.
[0006] Secondly, this application provides a method for image matting, the method comprising: Obtain the video to be cut out, and determine each video frame in the video to be cut out as the image to be cut out; Using the method for determining matting parameters described in the first aspect, the target matting parameters corresponding to each video frame of the video to be matted are determined; The video frames are rendered by matting according to the target matting parameters to generate the matted result image.
[0007] Thirdly, this application also provides an electronic device, including: a processor, a memory, and computer program instructions stored in the memory and executable on the processor; the processor, when executing the computer program instructions, implements the method as described in any one of the first to second aspects.
[0008] Fourthly, this application provides a computer-readable storage medium storing computer-executable instructions, which, when executed by a processor, are used to implement the method described in any one of the first to second aspects.
[0009] Fifthly, embodiments of this application provide a computer program product, including a computer program that, when executed by a processor, implements the method as described in any one of the first to second aspects.
[0010] Compared with the prior art, this application has the following advantages: The method for determining matting parameters provided in this application involves acquiring an image to be matted, determining ground truth segmentation information for each pixel in the image based on a pre-trained segmentation model. Since the ground truth segmentation information indicates whether a pixel is background or foreground, it accurately reflects whether each pixel in the image to be matted is foreground or background. Then, the matting parameters to be optimized are determined, i.e., initial matting parameters to be optimized are determined. These matting parameters control the overall matting intensity of the image. Next, the image to be matted is rendered according to the determined matting parameters to obtain effect segmentation information. Then, with the optimization goal of reducing the loss between the effect segmentation information and the ground truth segmentation information, the matting parameters are adjusted to finally obtain target matting parameters suitable for the current image to be matted. The entire process requires no manual parameter adjustment by the user and can automatically complete the optimization and determination of matting parameters. This improves the efficiency of matting parameter determination and ensures the accuracy of parameter determination, allowing users with different experience levels to consistently obtain suitable matting parameters and guaranteeing the final matting effect.
[0011] Furthermore, this application continuously optimizes and adjusts the matting parameters based on ground truth segmentation information when determining the matting parameters. This ensures that the final target matting parameters are adapted to the actual pixel distribution of the image to be matted. Compared to fixed or empirical parameters, this approach better adapts to matting scenarios with different color distributions, effectively avoiding issues such as excessive background residue or accidental removal of the foreground, thus significantly improving matting quality. Moreover, for video matting scenarios, this application can uniformly determine the target matting parameters for the entire video to be matted. This not only ensures the consistency of matting effects across different video frames and avoids flickering between frames, but also reduces the computational load associated with determining parameters frame by frame, improving the efficiency of video matting parameter determination. Attached Figure Description
[0012] Figure 1 This is a schematic diagram illustrating the application scenario of the solution provided in this application.
[0013] Figure 2 This is a flowchart illustrating an example of the method for determining matting parameters provided in the embodiments of this application.
[0014] Figure 3 This is a schematic diagram comparing the effects of image matting based on truth value segmentation information and image matting using related techniques in the embodiments of this application.
[0015] Figure 4 This is a schematic diagram comparing the cutout effect after adjusting the parameters in the second stage with the cutout effect after adjusting the parameters in the first stage in this application embodiment. The second stage parameter adjustment is to adjust in the direction of increasing the cutout intensity.
[0016] Figure 5 This is a schematic diagram comparing the cutout effect after adjusting the parameters in the second stage with the cutout effect after adjusting the parameters in the first stage in this embodiment of the application. The second stage parameter adjustment is to adjust in the direction of reducing the cutout intensity.
[0017] Figure 6 This is a flowchart illustrating another example of the method for determining matting parameters provided in the embodiments of this application.
[0018] Figure 7 This is a structural block diagram of the electronic device provided in this application. Detailed Implementation
[0019] To enable those skilled in the art to better understand the technical solutions of this application, the application will be clearly and completely described below with reference to the accompanying drawings of the embodiments. However, this application can be implemented in many other ways different from those described below. Therefore, based on the embodiments provided in this application, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of this application.
[0020] It should be noted that the terms "first," "second," "third," etc., in the claims, specification, and drawings of this application are used to distinguish similar objects and are not used to describe a specific order or sequence. Such data are interchangeable where appropriate so that the embodiments of this application described herein can be implemented in a sequence other than that shown or described in this application. Furthermore, the terms "comprising," "having," and their variations are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that includes a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to these processes, methods, products, or apparatuses.
[0021] It should be understood that in the embodiments of this application, "at least one" means one or more, and "more than one" means two or more. "And / or" is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, and B existing alone. The character " / " generally indicates that the related objects before and after it are in an "or" relationship. "Contains A, B and / or C" means containing any one, two, or three of A, B, and C.
[0022] It should be understood that in the embodiments of this application, "B corresponding to A", "B corresponding to A", "A corresponds to B" or "B corresponds to A" means that B is associated with A, and B can be determined based on A. Determining B based on A does not mean that B is determined solely based on A; B can also be determined based on A and / or other information.
[0023] To facilitate understanding of the various embodiments of this application, the application background of the embodiments will be explained.
[0024] With the rapid development of image and video technologies, the demand for image processing is increasing daily. Background removal, as a crucial operation, is widely used in various scenarios such as image editing, content creation, and green screen background removal for live streaming. The quality of background removal is highly dependent on the parameters used. These parameters control the leniency of background color judgment and determine the overall strength of the background removal. If the parameters are not set appropriately, problems such as the foreground being misjudged as background or excessive background residue can easily occur, affecting the final quality of the background removal. Traditionally, determining background removal parameters largely relies on manual user adjustments. Users need to repeatedly adjust parameter values based on the preview until the visual effect is satisfactory. This method is not only cumbersome and inefficient, but also difficult to consistently obtain optimal parameters due to differences in user visual experience, making it hard to guarantee consistent background removal quality. For users without professional image processing experience, it is even more difficult to quickly find suitable parameters, resulting in a poor user experience. Therefore, how to automatically determine appropriate background removal parameters and improve the efficiency and accuracy of parameter determination has become an urgent problem to be solved.
[0025] To address the above issues, embodiments of this application provide a method, apparatus, electronic device, and computer-readable storage medium for determining image matting parameters. This method automatically determines suitable matting parameters, effectively improving the efficiency and accuracy of parameter determination. It eliminates the need for repeated manual adjustments by the user, lowering the barrier to entry for parameter adjustment. Furthermore, it provides a more stable output of matting parameters adapted to the current image to be matted, resulting in a more stable matting effect and improved matting quality and user experience.
[0026] The method for determining image matting parameters provided in this application can be applied to image matting methods in various professional fields. Specifically, it can be applied to different scenarios such as green screen matting, portrait matting, and product image matting. It can process single static images to be matted or batch process video frames. This method can be integrated into various image processing software, live streaming software, and content creation tools, and executed when the corresponding program is run on an electronic device. The electronic device can be a personal computer, mobile device, server, or other different types of devices, which are not limited in this embodiment. Among them, green screen matting is an image segmentation technique based on color differences. It typically uses a green or blue background and removes the foreground subject from the green or blue background in a video or image using a specific algorithm, while retaining the foreground subject, thereby achieving the effect of changing the background or performing post-production compositing. It is widely used in film and television production, live streaming e-commerce, online education, and other scenarios.
[0027] To facilitate understanding of the method embodiments of this application, their application scenarios are described. Please refer to... Figure 1 , Figure 1This is a schematic diagram illustrating an application scenario of the solution provided in the embodiments of this application. This application scenario is merely an illustrative example and is not intended to limit the specific application scenario. Figure 1 As shown, in this application scenario, a server 102 and a client 101 are provided. In this embodiment, the client 101 and the server 102 establish a connection through network communication to transmit data.
[0028] Server 102 possesses high computing power. Server 102 can be a server, featuring high-speed central processing unit (CPU) computing power, long-term reliable operation, powerful input / output (I / O) external data throughput, and better scalability. Server 102 can be a single server or a server cluster. Server 102 stores the segmentation model provided in this application, used to complete the matting parameter determination process, and then returns the determined target matting parameters to the client. Server 102 can also provide other specific services to client 101, such as user information access, website access, application access, etc., which are not specifically limited in this application.
[0029] Client 101 can be an electronic device with display and data processing capabilities, such as a mobile phone, tablet, smartwatch, desktop computer, smart TV, VR device, in-vehicle device, wearable device, or laptop. Client 101 is used to send the image to be cut out to server 102 and receive the target cutout parameters sent by server 102. Client 101 can also be used to send access requests and interactive information to server 102, so that server 102 can send the corresponding request data to client 101 for display.
[0030] The client 101 and the server 102 can communicate using various communication systems, such as wired or wireless communication systems.
[0031] In other application scenarios, only client 101 may be included. Client 101 is configured with the segmentation model described in this application and a program for determining the matting parameters. The matting parameters are determined directly on the client side without interaction with the server. This method is more convenient for offline processing and can also improve the speed of determining the matting parameters, making it suitable for local edge image processing scenarios. The solution provided in this application can also be applied to other application scenarios, and this application is not specifically limited to them.
[0032] Example 1 The first embodiment of this application provides a method for determining matting parameters. This method can be applied to electronic devices, such as servers, desktop computers, laptops, mobile phones, tablets, smartwatches, smart TVs, VR devices, in-vehicle devices, wearable devices, and other electronic devices with data processing capabilities. Since determining the matting parameters may consume significant resources, the parameter determination process can typically be performed by a server.
[0033] like Figure 2 As shown, the method for determining the matting parameters provided in the first embodiment of this application includes the following steps S110 to S150.
[0034] Step S110: Obtain the image to be cut out.
[0035] The image to be cut out can be any image containing the foreground to be extracted, or a video frame in a video, such as a green screen image in a green screen live broadcast scene, a portrait photo whose background needs to be changed, a product display image that needs to be edited, a set of video frames that need to be cut out, etc. This application does not impose any restrictions on this.
[0036] It can directly read the image to be cut out from the local storage of the electronic device, or receive the image to be cut out sent by other devices, or determine the image to be cut out as the user-uploaded image, depending on the actual application scenario.
[0037] Step S120: Determine the ground truth segmentation information corresponding to each pixel in the image to be cut out based on the pre-trained segmentation model. The ground truth segmentation information is used to indicate whether the corresponding pixel is the background or the foreground.
[0038] When the image to be cut out obtained in step S110 is an image, the segmentation model will output the corresponding ground truth segmentation information for each pixel of the image, and the judgment can be completed directly pixel by pixel.
[0039] If the obtained data consists of multiple consecutive video frames, keyframes in the video can be selected, or a preset number of video frames can be uniformly extracted and input into the segmentation model. Alternatively, each video frame can be input into the segmentation model sequentially to obtain the ground truth segmentation information corresponding to each video frame input into the segmentation model.
[0040] The pre-trained segmentation model described above possesses semantic segmentation capabilities, accurately distinguishing foreground and background pixels in the image to be matted. The ground truth segmentation information output can serve as a standard reference for subsequent matting parameter adjustments. The specific structure of the segmentation model can be selected according to requirements. For example, it can be a semantic segmentation network built based on deep learning convolutional neural networks, a deep neural network based on a nested U-shaped structure, or other types of models with pixel-level segmentation capabilities. This application does not impose specific limitations. In particular, each feature extraction module in the encoder and decoder of the deep neural network based on a nested U-shaped structure is composed of a Residual Symmetry U-Net (RSU) module, which can extract the contextual information of the image at different scales, taking into account both local detail features and global semantic information, thereby improving the accuracy of the segmentation results and ensuring that the reference standard for subsequent matting parameter optimization is accurate and reliable.
[0041] In one implementation, the above segmentation model is trained through the following steps A to C.
[0042] Step A: Obtain training samples, which include sample images and corresponding ground truth segmentation information. The sample images are HSV color space images.
[0043] The ground truth segmentation information of the samples is used to indicate whether each pixel in the sample image is a foreground or background pixel. It is obtained through manual annotation or generated using a high-precision image matting tool, and serves as the annotation benchmark for training the segmentation model. In the HSV color space, the color information of the image makes it easier for the model to learn to distinguish between foreground and background, improving the efficiency of segmentation model training and the final segmentation accuracy. The HSV color space represents pixel color through three dimensions: hue, saturation, and lightness, which is closer to how the human eye perceives color and is more conducive to distinguishing the color differences between background and foreground.
[0044] In this embodiment, when the original sample image is an RGB or other non-HSV color space image, the original sample image can be converted to the HSV color space to obtain the sample image.
[0045] Step B: Apply random perturbation to at least one of the hue channel, saturation channel, and brightness channel of the sample image to generate a sample enhancement image.
[0046] Specifically, random perturbations can be applied to one of the hue, saturation, and brightness channels, or random perturbations can be applied to two or three channels simultaneously. The range of random perturbations can be preset according to training requirements. For example, the perturbation range of the hue channel can be set to [-X, X], and the perturbation ranges of the saturation and brightness channels can be set to [Y1, Y2] respectively. After generating independent random perturbation values for each channel according to the corresponding random perturbation range, these values are superimposed on the pixel values of the original channel to obtain the enhanced sample image.
[0047] By applying random perturbations, the color distribution diversity of training samples can be enriched, improving the model's robustness to foreground and background segmentation under different lighting conditions and shooting color differences, and avoiding the problem of poor model generalization ability caused by the monotonous color distribution of training samples. For example, by applying random offset perturbations to the hue channel, the influence of different color temperature light sources and white balance differences on image hue can be simulated; by applying random scaling to the saturation channel, the influence of different ambient humidity and light intensity on image saturation can be simulated; and by applying random increases or decreases to the brightness channel, the image brightness fluctuations caused by changes in fill light brightness and alternating ambient light intensity in live streaming scenes can be simulated.
[0048] By perturbing the HSV color space instead of the RGB color space, the coupling between the three perturbation dimensions is lower, resulting in a more natural and controllable enhancement effect that more closely resembles the real-world distribution of lighting and device differences. This data augmentation strategy significantly improves the effective diversity of the training data, helping to prevent the model from overfitting to specific lighting or device parameters in the training set and improving the accuracy of image matting using ground truth masks. Figure 3 As shown, for Figure 3 When performing image matting on the image to be matted shown in (a), the segmentation result obtained according to relevant techniques is as follows: Figure 3 As shown in (b) above, the segmentation result obtained according to the segmentation model provided in this application is as follows: Figure 3 As shown in (c), the comparison shows that the top of the image in the related technology has a serious background omission, that is, a large area of the background region is misjudged as the foreground. However, the segmentation model of this application can accurately identify the background region and obtain complete and accurate segmentation results. It can be seen that the segmentation model trained by this application has better segmentation accuracy.
[0049] Step C: Train the segmentation model based on the sample image, the sample augmented image, and the corresponding sample ground truth segmentation information.
[0050] Specifically, the original sample images and the generated enhanced sample images can be mixed to form a training dataset. The corresponding ground truth segmentation information is used as a supervision signal and substituted into the model for end-to-end training. The model parameters are updated through backpropagation until the model's segmentation accuracy on the validation set reaches a preset requirement or the number of training iterations reaches a preset limit, at which point training stops and the segmentation model is obtained. By using the method of training with both original and enhanced samples, the model can learn the distinguishing features between foreground and background under different colors and lighting conditions, further improving the model's generalization ability and segmentation accuracy.
[0051] In one specific embodiment, the ground truth segmentation information corresponding to each pixel in the image to be cut out can be a ground truth mask corresponding to the image to be cut out. The ground truth mask has the same size as the image to be cut out, and the value of each pixel position corresponds to the binary classification result of the pixel as foreground or background. For example, each pixel with a value of 0 represents the background, and a value of 1 represents the foreground. Other representation methods can also be set according to requirements, and this application does not limit them. In particular, storing the ground truth segmentation information in the form of a mask can more conveniently identify the foreground and background information of each pixel, and also facilitates the subsequent calculation of the loss between the effect segmentation information and the ground truth segmentation information, thereby improving the computational efficiency.
[0052] Step S130: Determine the matting parameters to be optimized, which are used to control the overall matting intensity of the image.
[0053] In this step, the initially determined matting parameters to be optimized can be preset default values; alternatively, an initial matting parameter to be optimized can be estimated based on the background type and image type of the image to be matted. For example, for scenes with a uniform background color distribution, such as a pure green green screen, the initial matting parameters can be set to a relatively moderate default value, such as any value between 0.4 and 0.6 for the similarity parameter. For scenes with a certain color deviation in the background, the initial parameter value can be appropriately increased to accelerate the convergence speed of subsequent parameter optimization. This initial parameter is the object to be optimized, and subsequent iterations will gradually yield more suitable target matting parameters; alternatively, the initial matting parameters to be optimized can be randomly generated. Different initialization methods do not affect the accuracy of the final optimized target matting parameters.
[0054] Step S140: Perform cutout rendering on the image to be cut out according to the cutout parameters to obtain the effect segmentation information corresponding to each pixel in the image to be cut out.
[0055] Performing matting rendering on the image to be matted according to the matting parameters refers to calculating the determination result of whether each pixel is judged as foreground or background based on the current matting parameters and a preset matting algorithm, thereby obtaining the effect segmentation information corresponding to each pixel. The effect segmentation information is identified by the classification result of the corresponding pixel, maintaining consistency with the representation method of the ground truth segmentation information to facilitate subsequent loss calculation and comparison. For example, the effect segmentation information can be represented in the form of a mask, where the value of each pixel corresponds to the classification result of the pixel as foreground or background under the current matting parameters, facilitating subsequent comparison calculation with the ground truth segmentation information.
[0056] Different matting algorithms use matting parameters differently. Taking green screen matting as an example, the color similarity between a pixel and a preset background color is typically compared with the matting parameters. If the similarity is greater than the matting parameters, it is determined to be the background; if the similarity is less than the matting parameters, it is determined to be the foreground. It is evident that the matting parameters determine the threshold range for background color determination, thus affecting the final matting result. Different determination logics can be executed according to the rules of the corresponding matting algorithm; this application does not impose specific limitations.
[0057] Step S150: With the optimization objective of reducing the loss between the effect segmentation information and the ground truth segmentation information, adjust the matting parameters to obtain the target matting parameters.
[0058] Specifically, the matting parameters can be iteratively optimized using gradient descent. After each iteration, the loss between the effective segmentation information and the ground truth segmentation information is recalculated, and the values of the matting parameters are adjusted according to the gradient of the loss until the loss meets the preset convergence condition or reaches the preset maximum number of iterations. Then, the iteration stops, and the final parameters are determined as the target matting parameters. The loss can employ common loss calculation methods such as cross-entropy loss and mean squared error loss. The specific loss calculation method can be selected according to the representation method of the segmentation information; this application does not impose specific limitations.
[0059] In this embodiment, when the image to be matted is a single image, the above parameter optimization process can be directly performed on this image to obtain the target matting parameters adapted to the image. When the image to be matted is a video containing multiple video frames, the above parameter optimization process can be performed on each video frame separately to obtain the target matting parameters corresponding to each video frame. This can adapt to changes in background light and color between different frames, ensuring that each frame can obtain a good matting effect. Alternatively, multiple video frames can be extracted, the total loss of all extracted frames can be calculated, and the same matting parameter can be adjusted based on the total loss to obtain the target matting parameters that are universal for the entire video. This can reduce the amount of parameter calculation, improve the speed of parameter determination, and is suitable for scenarios where the overall background conditions of the video are relatively stable, and the target matting parameters corresponding to each video frame are the same. Alternatively, the corresponding parameters for each video frame can be... The target matting parameters are smoothed, for example, by interpolating parameters between adjacent frames to avoid abrupt changes in parameters that could cause significant jumps in the matting effect, thus improving the smoothness of the video matting result. Alternatively, the calculated target matting parameters for each video frame can be averaged to obtain universal target matting parameters for the entire video, which also simplifies the calculation process and adapts to video matting scenarios with stable backgrounds, ensuring that the final target matting parameters for each video frame are the same. Another approach is to determine the total loss between the effect segmentation information and the ground truth segmentation information based on the effect segmentation information for each video frame, and then uniformly optimize to obtain global target matting parameters suitable for the entire video. Specifically, the appropriate processing method can be selected based on the differences in foreground and background distribution between video frames; this application does not specifically limit this. When the target matting parameters for each video frame are the same, frequent parameter changes can prevent inter-frame flickering in the matting result, improving the visual stability of the video matting.
[0060] This application mainly uses the example of a single image to be cut out. When the image to be cut out includes multiple video frames, the parameter determination process is similar to that when the image to be cut out is a single image. Those skilled in the art can make adaptive adjustments based on the above description, and this application will not elaborate further.
[0061] In one implementation, step S150 can be implemented according to the following steps S151 to S152.
[0062] Step S151: With the optimization goal of reducing the pixel-level difference between the effect segmentation information and the ground truth segmentation information, the matting parameters are adjusted in the first stage to obtain the first stage parameters.
[0063] The aforementioned pixel-level difference refers to the difference in classification results between the effect segmentation information and the ground truth segmentation information for each pixel. Using pixel-level difference as the optimization target in the first stage can quickly bring the matting parameters to a roughly suitable value range, accelerating the overall optimization speed. Since the parameters are adjusted directly based on the classification differences of each pixel, this stage can quickly correct large deviations between the initial parameters and the target parameters, laying a solid foundation for further adjustments.
[0064] Specifically, the classification difference between the effect segmentation information and the ground truth segmentation information of each pixel can be calculated. The differences of each pixel are summarized to obtain the overall pixel-level total difference. Then, the matting parameters are updated in a gradient descent manner based on the pixel-level total difference. The iterative process of updating the matting parameters, recalculating the difference, and adjusting the parameters is repeated until the overall pixel-level difference drops to the first stage preset threshold, or the preset number of first stage iterations is completed. Then, the first stage iteration is stopped, and the current parameters are determined as the first stage parameters.
[0065] In one specific embodiment, step S151 can be implemented according to the following steps S151a~S151c.
[0066] Step S151a: Based on the segmentation region corresponding to each pixel in the true value segmentation information, determine the loss weight corresponding to each pixel, wherein the loss weight corresponding to the background segmentation region is greater than the loss weight corresponding to the foreground segmentation region.
[0067] The aforementioned segmented regions may include background segmented regions, foreground segmented regions, or other regions, such as uncertain regions. This embodiment uses the division of foreground and background segmented regions for illustration.
[0068] In real-world background removal scenarios, such as green screen background removal, the misclassification of the background area as the foreground has a greater impact on the final background removal effect than the minor misclassification within the foreground. This can result in a large amount of background remaining around the foreground, severely affecting the visual quality of the background removal. Therefore, assigning a greater loss weight to the background segmentation area allows the optimization process to focus more on the misclassification of the background area, prioritize correcting background misclassification, and improve the visual quality of the final background removal result.
[0069] Specifically, the loss weight corresponding to pixels belonging to the background segmentation region can be a preset multiple of the pixels belonging to the background segmentation region. This preset multiple is greater than 1, and can be set to any value between 2 and 10. The specific weight can be adjusted according to actual needs, and this application does not impose any specific limitations. Weighting amplifies the loss caused by missegmentation of the background region, guiding parameter optimization towards reducing background missegmentation, thus better meeting the visual effect requirements of actual image matting.
[0070] In one specific embodiment, step S151a, when determining the loss weight corresponding to each pixel, can also obtain a preset cutout region; set the loss weight of each pixel within the cutout region to 0; and determine the loss weight of each pixel in the region outside the cutout region in the image to be cut out based on the segmentation region corresponding to each pixel in the ground truth segmentation information. In other words, if the user pre-labels the cutout region in the image that does not need processing, the user may directly classify this region as foreground or background. Therefore, misclassification of pixels in this region does not affect the final cutout effect. Setting the loss weight of this region to 0 can avoid the misclassification result affecting the overall parameter optimization, allowing the optimization process to focus more on the region that the user needs to process, further improving the accuracy of the target cutout parameters, reducing computational load, and increasing optimization speed.
[0071] Step S151b: The pixel-level difference between the effect segmentation information and the ground truth segmentation information is weighted according to the loss weight corresponding to each pixel to obtain the pixel-level weighted error.
[0072] Specifically, the classification difference between the effect segmentation information and the ground truth segmentation information can be calculated for each pixel, multiplied by the loss weight corresponding to that pixel, and then summed to obtain the final pixel-level weighted error. The weighted error can reflect a higher penalty for misclassification of the background area. Subsequent parameter adjustments based on this error can make the parameter optimization direction more in line with the visual needs of actual image matting.
[0073] The above pixel-level weighted error can also be understood as pixel-level loss, which can be expressed by the following formula (1).
[0074] (1) Where s represents the matting parameters to be optimized, i,j represents the pixel coordinates at the i-th row and j-th column position in the image, W(i,j) represents the loss weight corresponding to the pixel at the i-th row and j-th column position, and M... s (i,j) represents the value of the segmentation information at that pixel position under the current matting parameters s, M (i,j) represents the value of the truth segmentation information at that pixel position. This formula calculates... The normalized pixel-level weighted error accurately reflects the weighted difference between the effect segmentation and the ground truth segmentation under the current matting parameters.
[0075] The loss weight corresponding to each pixel can be represented by the mask weight matrix W. For example, the mask weight matrix W can be represented by the following formula (2).
[0076] (2) in, This means that the pixel belongs to the area that does not need to be cut out. This means that the pixel belongs to the background area. This indicates that the pixel is the foreground region. In formula (2), the loss weight corresponding to the background region is set to 5, which is greater than the weight of 1 for the foreground region. This meets the aforementioned optimization requirement of giving higher penalties to background misclassification. The specific weight value can be adjusted according to the actual scenario. This example does not limit it.
[0077] Step S151c: With minimizing the pixel-level weighted error as the optimization objective, the matting parameters are adjusted in the first stage to obtain the first stage parameters.
[0078] Specifically, the gradient of the matting parameters can be calculated based on the current pixel-level weighted error, and then the values of the matting parameters can be updated in the opposite direction of the gradient to complete one parameter adjustment. After that, the pixel-level weighted error is recalculated based on the updated matting parameters, and the iterative steps of gradient calculation and parameter update are repeated until the pixel-level weighted error meets the convergence condition of the first stage or reaches the maximum number of iterations preset in the first stage. Then, the first stage iteration is stopped, and the parameters at this time are determined as the parameters of the first stage.
[0079] In this embodiment, when determining the parameters for the first stage, the loss weight is determined by the segmentation region to which each pixel belongs. A higher loss weight is set for the background segmentation region, which allows the parameter adjustment in the first stage to be more biased towards correcting the missegmentation problem of the background region, reducing the impact of background residue on the matting effect, and quickly converging the matting parameters to a reasonable value range, providing a good foundation for subsequent optimization.
[0080] Step S152: With the optimization objective of reducing the structural difference between the effect segmentation information and the ground truth segmentation information, the parameters of the first stage are adjusted in the second stage to obtain the target matting parameters.
[0081] The aforementioned structural-level differences refer to the differences between the effect segmentation information and the ground truth segmentation information in terms of the overall structure, such as the edge structure and regional distribution of the foreground and background. Compared to pixel-level differences, which focus on the correctness of classification of individual pixels, structural-level differences focus on the overall structure of the matting result and the degree of matching between it and the ground truth segmentation. This allows for further optimization of the segmentation effect after the parameters in the first stage have converged to a reasonable range, resulting in a more natural overall structure that better matches the realistic foreground contours. Building upon the correction of large deviations and resolution of background residue issues in the first stage, optimizing structural consistency in the second stage further enhances the overall visual quality of the matting result, yielding parameter results that more closely match the user-annotated ground truth values.
[0082] Specifically, the structural loss between the effect segmentation information and the ground truth segmentation information can be calculated, and the parameters of the first stage can be iteratively adjusted using the gradient descent method. The structural loss is recalculated after each iteration until the structural loss meets the convergence condition of the second stage or reaches the maximum number of iterations preset in the second stage. Then, the iteration stops, and the final parameters are determined as the target matting parameters.
[0083] The structural loss can be calculated using loss functions applicable to assessing differences in segmented structures in related technologies, such as IoU loss, structural similarity loss, etc. This application does not make any specific limitations.
[0084] In one implementation, step S152 can be implemented according to the following steps S152a~S152b.
[0085] Step S152a: When there is a first connected component in the background region of the truth segmentation information that is determined to be the foreground by the effect segmentation information, the parameters of the first stage are adjusted in the second stage in the direction of increasing matting intensity to reduce the number of the first connected components.
[0086] The aforementioned first connected region refers to a connected pixel block in the background region that has been misclassified as foreground. A first connected region is formed by the aggregation of multiple consecutive background residual pixels. The more first connected regions there are and the larger their area, the more serious the background residue problem is.
[0087] When the background region of the truth segmentation information contains a first connected component that is identified as the foreground by the effect segmentation information, the following will occur: Figure 4 The phenomenon shown in (a) is the residual background impurities after image matting based on effect segmentation information, i.e., the background is missed. These residual impurities, which should have been removed as background, are mistakenly retained as foreground. These scattered small connected components affect the appearance of the matting result and can also cause unnecessary noise in the subsequent compositing. Therefore, when such first connected components are detected, the parameters can be adjusted in the direction of increasing matting intensity. This can improve the strictness of background judgment, allowing the originally misclassified background areas to be re-judged as background, thereby reducing the number of residual background impurities and improving the cleanliness of the matting result. After the second stage of adjusting the parameters in the direction of increasing matting intensity, the matting effect is as follows: Figure 4 As shown in (b) of the diagram.
[0088] When there is a first connected component in the background region of the ground truth segmentation information that is determined to be the foreground by the effect segmentation information, it indicates that there is a structural difference between the effect segmentation information and the ground truth segmentation information. This structural difference is mainly reflected in the fact that the connected structure distribution of the missegmented region does not match the ground truth. Therefore, adjusting the parameters based on this structural difference can specifically solve the structural problem of the background residue and make the structure of the segmentation result more consistent with the ground truth.
[0089] When adjusting the parameters of the first stage in the second stage in the direction of increasing image matting intensity, the adjustment step size can be determined based on the number and area of the first connected regions. The more numerous and larger the total area of the first connected regions, the larger the parameter adjustment step size and the greater the adjustment range, enabling rapid correction of large-scale background residue. If only a few small-area first connected regions exist, a smaller adjustment step size is used to avoid excessive adjustment range leading to over-mutilation of the foreground edges. When the number of first connected regions is less than a preset threshold, or the total area of the first connected regions is less than a preset area threshold, the adjustment can be stopped, indicating that the structural error of the current parameters has met the requirements.
[0090] When the matting parameter is the aforementioned similarity parameter, since the larger the similarity parameter value, the stronger the matting intensity, and the easier it is to determine the pixel as the background, the second stage parameter adjustment of the first stage parameter can be performed in the direction of increasing the matting intensity, which can be done in the direction of increasing the similarity parameter; if the matting parameter is of other types, it can also be adjusted according to its correlation with the matting intensity, and this application does not make specific limitations.
[0091] In one specific embodiment, when adjusting the parameters of the first stage in the direction of increasing matting intensity in the second stage, the adjustment can be made with the limitation that the number of third connected regions does not exceed a preset threshold. The third connected region is a connected region within the foreground region of the ground truth segmentation information that is determined as background by the effect segmentation information and intersects with the boundary of the foreground region of the ground truth segmentation information.
[0092] The aforementioned third connected region can also be understood as a boundary erosion type connected region, that is, a connected pixel block that appears on the boundary of the ground truth foreground region but is misclassified as background. This type of connected region causes the foreground contour edge to be unnecessarily removed, resulting in missing foreground edges and affecting the integrity of the contour. The boundary erosion effect caused by the third connected region is as follows: Figure 5 As shown in (a) of the diagram.
[0093] This embodiment limits the number of third connected regions to a preset threshold, which avoids excessively increasing the matting intensity to eliminate background residue, resulting in excessive erosion of the foreground edges. It protects the integrity of the foreground edges while eliminating background residue, balancing the needs of background residue correction and foreground contour integrity, and obtaining a more reasonable segmentation result.
[0094] Step S152b: When there is a second connected component in the foreground region of the truth segmentation information that is determined to be the background by the effect segmentation information, and there is no first connected component in the background region of the truth segmentation information, the parameters of the first stage are adjusted in the second stage in the direction of decreasing matting intensity to reduce the number of the second connected components.
[0095] The aforementioned second connected region refers to a connected pixel block in the foreground region that has been misclassified as background. A second connected region is formed by a cluster of multiple consecutive pixels that have been cut out in the foreground. The more second connected regions there are and the larger their area, the more serious the foreground missing problem is.
[0096] When a second connected component, identified as background by the effect segmentation information, exists within the foreground region of the truth segmentation information, it can lead to partial foreground loss after masking using the effect segmentation information. These missing foreground components, which should have been preserved, are mistakenly identified as background and removed, thus damaging the integrity of the foreground outline and affecting the visual effect of the masking result. The second connected component is the mis-masked foreground region. Figure 5 As shown in (a) above. When there is no remaining background in the first connected region, but a second connected region exists, it indicates that the current matting intensity is too high. Therefore, adjusting the parameters in the direction of decreasing matting intensity can reduce the strictness of background judgment, allowing the originally misclassified foreground regions to be reclassified as foreground, thereby reducing the situation of missing foreground and obtaining a complete foreground segmentation result. The matting effect after adjusting the parameters in the direction of decreasing matting intensity is shown below. Figure 5 As shown in (b) above. Similarly, the adjustment step size can also be determined based on the total number and total area of the second connected components. The more the total number and the larger the total area of the second connected components, the larger the adjustment step size can be set. When the number of second connected components is less than the preset number threshold, or the total area is less than the preset area threshold, the adjustment can be stopped, and the adjusted parameters can be determined as the target matting parameters.
[0097] When the matting parameter is the aforementioned similarity parameter, when adjusting the first stage parameter in the second stage in the direction of decreasing matting intensity, the adjustment can be made in the direction of decreasing similarity parameter; if the matting parameter is of other types, it can be adjusted according to its correlation with matting intensity.
[0098] When adjusting the parameters of the first stage in the second stage in the direction of decreasing matting intensity, the first stage parameters can be adjusted in the direction of decreasing matting intensity, with the constraint that the image structure similarity between the effect segmentation information and the ground truth segmentation information is not lower than a preset similarity threshold.
[0099] Among them, image structure similarity is used to represent the degree of similarity between the effect segmentation information and the ground truth segmentation information in the overall structure. When it is not lower than the preset similarity threshold, it can ensure that the segmentation structure will not deviate too much from the ground truth structure after parameter adjustment, and avoid introducing new background residue problems due to reducing the matting intensity. Using the similarity of the two in the overall structure as a background structure consistency constraint can ensure the overall structure quality and background cleanliness while repairing foreground loss, and avoid new misclassification problems caused by excessive parameter adjustment.
[0100] Specifically, the similarity in overall structure can be calculated using the intersection-union ratio (IUU) of the background regions in the effect segmentation information and the background regions in the ground truth segmentation information. A higher IUU indicates a higher structural similarity between the two. Alternatively, the overall structural similarity can be calculated based on the structural similarity index of the overall segmentation masks corresponding to the effect segmentation information and the ground truth segmentation information, respectively. Other methods can also be used to calculate image structural similarity, and this application does not impose any specific limitations on these methods.
[0101] In one specific embodiment, the image structural similarity described above is determined through the following steps a~d.
[0102] Step a: Determine the standard structural similarity between the effect segmentation information and the truth segmentation information.
[0103] The aforementioned standard structural similarity is used to measure the similarity between two images. It is used to comprehensively evaluate the brightness, contrast, and structural information of two images by simulating the characteristics of the human visual system, so as to obtain a structural similarity evaluation result that is more in line with the subjective perception of the human eye and has high evaluation accuracy.
[0104] The aforementioned standard structural similarity can be the normalized similarity calculated based on the overall structural features between the effect segmentation mask corresponding to the current effect segmentation information to be evaluated and the ground value segmentation mask corresponding to the ground value segmentation information. The value range is usually from 0 to 1. The larger the value, the higher the degree of structural matching between the two, and the smaller the value, the greater the structural difference.
[0105] Step b: Extract the edge contours of the effect segmentation information and the ground truth segmentation information respectively to obtain the first edge map and the second edge map.
[0106] Specifically, edge detection algorithms can be used to extract edges from the effect segmentation mask corresponding to the effect segmentation information and the ground value segmentation mask corresponding to the ground value segmentation information, respectively obtaining the first edge map of the effect segmentation and the second edge map of the ground value segmentation. Edge extraction can focus on the contour structure differences of the segmentation results, highlighting the foreground edge, a structural feature that has a significant impact on the matting effect, making the structural similarity assessment more in line with the needs of the matting scenario.
[0107] Step c: Determine the edge structure similarity between the first edge map and the second edge map.
[0108] Specifically, the degree of edge matching at corresponding positions in the first and second edge maps can be used as the edge structure similarity; the higher the edge overlap, the higher the edge structure similarity. Alternatively, the edge structure similarity can be obtained by calculating the intersection-union ratio (IUU) of the two edge maps or by calculating the matching ratio of edge pixels; this application does not impose any limitations on this.
[0109] Step d: Determine the image structure similarity between the effect segmentation information and the ground truth segmentation information based on the standard structure similarity and the edge structure similarity.
[0110] Specifically, the weighted average of the standard structural similarity and the edge structural similarity can be used as the image structural similarity. Weights can be assigned to the standard structural similarity and the edge structural similarity, respectively. For example, the weight ratio can be adjusted according to the contour accuracy requirements of the matting scenario. For instance, in scenarios with high contour accuracy requirements, the weight of the edge structural similarity can be increased, allowing the structural similarity evaluation to focus more on the matching degree of the edge contours, which better meets the parameter optimization requirements of high-precision matting. Alternatively, the image structural similarity can be obtained by determining the sum of the standard structural similarity and the edge structural similarity, or through other combinations. This application does not impose specific limitations.
[0111] By combining the standard structural similarity and edge structural similarity of the overall image to calculate the final image structural similarity, we can take into account both the overall distribution consistency of the region and highlight the impact of the foreground contour edges on the matting effect. This will result in a structural similarity evaluation result that is more in line with the needs of the matting scenario, providing a more accurate constraint basis for parameter adjustment and ensuring that the structural differences after parameter adjustment meet the optimization target requirements.
[0112] This embodiment adjusts the matting parameters based on the existence state of connected components in different cases, which can specifically solve different types of structural misclassification problems. While ensuring a clean background, it preserves the foreground as completely as possible, making the structure of the final segmentation result more consistent with the ground truth segmentation and achieving better matting results.
[0113] In one specific embodiment, step d above can be implemented according to the following steps d1~d2.
[0114] Step d1: Determine the background coverage integrity value based on the number of first pixels in the background region of the true value segmentation information that are determined as the foreground region by the effect segmentation information, wherein the background coverage integrity value is negatively correlated with the number of first pixels.
[0115] Step d2: Using the background coverage integrity value as the weight, the standard structural similarity and the edge structural similarity are weighted to obtain the image structural similarity between the effect segmentation information and the ground truth segmentation information.
[0116] For example, the more pixels in the background region of the ground truth segmentation information are misclassified as foreground, the more serious the background misclassification problem is and the lower the background coverage integrity is. Therefore, the first pixel count and the background coverage integrity value are set to be negatively correlated. The more first pixels there are, the smaller the background coverage integrity value is. In this case, when the image structure similarity is calculated using the background coverage integrity value as a weighted weight, the overall structure similarity score can be reduced. This allows the image structure similarity to more accurately reflect the background residue problem in the current segmentation result, avoid misjudging the structure similarity as meeting the requirements when there is obvious background residue, and more accurately reflect the actual structure quality of the current segmentation result, providing more accurate constraints for parameter adjustment.
[0117] For example, the image structure similarity between the effect segmentation information and the ground truth segmentation information can be determined by the following formula (3).
[0118] (3) in, The image structure similarity between the effect segmentation information and the ground truth segmentation information. For the above standard structural similarity, For the above edge structure similarity, The proportion of the number of first pixels in the background region that are misclassified as foreground pixels to the total number of pixels in the background region, representing the true segmentation information. This indicates the background coverage integrity value mentioned above.
[0119] The constraint that the image structure similarity between the effect segmentation information and the ground truth segmentation information is not lower than a preset similarity threshold can be expressed by the following formula (4).
[0120] (4) in, This is a preset similarity threshold.
[0121] This implementation method optimizes parameters in two stages. The first stage quickly converges and corrects large deviations, prioritizing the resolution of background misclassification issues that have a greater impact. The second stage further optimizes structural consistency to improve the overall matting effect. This approach can improve the overall optimization speed and obtain target matting parameters that better meet visual requirements, thus balancing optimization efficiency and final matting effect.
[0122] In this embodiment of the application, when neither the first connected component in step S152a nor the second connected component in step S152b exists, it indicates that the structural misclassification problem has met the requirements after the current first-stage parameters have been adjusted, and the second-stage parameter adjustment can be stopped directly, and the first-stage parameters can be determined as the target matting parameters.
[0123] Once the target cutout parameters are determined, they can be provided to the user, allowing the user to directly use them for cutout processing without having to manually adjust the parameters repeatedly. This effectively reduces the user's adjustment costs and improves cutout efficiency and the quality of the final cutout effect.
[0124] The method for determining matting parameters provided in this application involves acquiring an image to be matted, determining ground truth segmentation information for each pixel in the image based on a pre-trained segmentation model. Since the ground truth segmentation information indicates whether a pixel is background or foreground, it accurately reflects whether each pixel in the image to be matted is foreground or background. Then, the matting parameters to be optimized are determined, i.e., initial matting parameters to be optimized are determined. These matting parameters control the overall matting intensity of the image. Next, the image to be matted is rendered according to the determined matting parameters to obtain effect segmentation information. Then, with the optimization goal of reducing the loss between the effect segmentation information and the ground truth segmentation information, the matting parameters are adjusted to finally obtain target matting parameters suitable for the current image to be matted. The entire process requires no manual parameter adjustment by the user and can automatically complete the optimization and determination of matting parameters. This improves the efficiency of matting parameter determination and ensures the accuracy of parameter determination, allowing users with different experience levels to consistently obtain suitable matting parameters and guaranteeing the final matting effect.
[0125] Furthermore, this application continuously optimizes and adjusts the matting parameters based on ground truth segmentation information when determining the matting parameters. This ensures that the final target matting parameters are adapted to the actual pixel distribution of the image to be matted. Compared to fixed or empirical parameters, this approach better adapts to matting scenarios with different color distributions, effectively avoiding issues such as excessive background residue or accidental removal of the foreground, thus significantly improving matting quality. Moreover, for video matting scenarios, this application can uniformly determine the target matting parameters for the entire video to be matted. This not only ensures the consistency of matting effects across different video frames and avoids flickering between frames, but also reduces the computational load associated with determining parameters frame by frame, improving the efficiency of video matting parameter determination.
[0126] The following describes the model training method provided in this application through an exemplary process, such as... Figure 6 As shown, the model training method in this example includes the following steps S1 to S5.
[0127] Step S1: Obtain the image to be cut out and the preset cutout area.
[0128] Step S2: Determine the ground truth mask corresponding to the image to be cut out based on the pre-trained segmentation model.
[0129] The truth mask is used to indicate whether each pixel in the image to be cut out belongs to the foreground or the background.
[0130] Step S3: Construct a mask weight matrix based on the cutout region and the ground truth mask.
[0131] Step S4: With the optimization goal of reducing the pixel-level difference between the effect mask and the ground truth mask, the parameters of the matting to be optimized are adjusted in the first stage to obtain the first stage parameters.
[0132] The effect mask is the mask obtained after rendering the image to be cut out according to the cutout parameters to be optimized.
[0133] Step S5: With the optimization goal of reducing the structural difference between the effect mask and the ground truth mask, the parameters of the first stage are adjusted in the second stage to obtain the target matting parameters.
[0134] In the second stage of parameter adjustment, if the structural differences do not meet the preset requirements, it indicates that there is a problem with the subjective effect. Then, step S5 is executed to adjust the parameters of the first stage in the second stage until the structural differences meet the preset requirements and the target matting parameters are obtained.
[0135] Example 2 The second embodiment of this application also provides a method for image matting, which is applied to electronic devices, such as servers, desktop computers, laptops, mobile phones, tablets, smartwatches, smart TVs, VR devices, in-vehicle devices, wearable devices, and other electronic devices with data processing functions. The image matting method includes the following steps S210 to S230.
[0136] Step S210: Obtain the video to be cut out.
[0137] Obtain the video to be cut out, and determine each video frame in the video to be cut out as the image to be cut out; Step S220: Using the method for determining matting parameters described in any one of the first embodiments, determine the target matting parameters corresponding to each video frame of the video to be matted.
[0138] The target matting parameters are the same for each video frame.
[0139] Step S230: Perform matting rendering on each of the video frames according to the target matting parameters to generate a matting result image.
[0140] The image matting method provided in this application can effectively improve the overall image quality while avoiding the generation of false details and textures that do not conform to the original image. The output results are more realistic and have a more natural visual experience.
[0141] Example 3 The third embodiment of this application also provides a device for determining matting parameters corresponding to the method embodiment for determining matting parameters provided in the first embodiment. Since the device embodiment is basically similar to the method embodiment, it is described simply. For details of the relevant technical features and their effects, please refer to the corresponding descriptions of the matting method embodiments provided above. The matting method device provided in this embodiment includes: Image acquisition unit, used to acquire the image to be cut out; The truth value determination unit is used to determine the truth value segmentation information corresponding to each pixel in the image to be cut out based on a pre-trained segmentation model. The truth value segmentation information is used to indicate whether the corresponding pixel is the background or the foreground. The first parameter determination unit is used to determine the matting parameters to be optimized, wherein the matting parameters are used to control the overall matting intensity of the image; The image cutout rendering unit is used to perform image cutout rendering on the image to be cut out according to the cutout parameters, so as to obtain the effect segmentation information corresponding to each pixel in the image to be cut out; The parameter optimization unit is used to adjust the matting parameters with the optimization objective of reducing the loss between the effect segmentation information and the ground truth segmentation information, so as to obtain the target matting parameters.
[0142] Example 4 The fourth embodiment of this application also provides an embodiment of an electronic device. The following description of the electronic device embodiment is merely illustrative. The electronic device embodiment is as follows: Please refer to Figure 7 Understanding the above electronic devices, Figure 7 This is a schematic diagram of an electronic device. The electronic device provided in this embodiment includes: a processor 1001, a memory 1002, a communication bus 1003, and a communication interface 1004; The memory 1002 is used to store computer instructions for data processing. When the computer instructions are read and executed by the processor 1001, they execute the method described in the first embodiment or the second embodiment.
[0143] Example 5 The fifth embodiment of this application also provides a computer-readable storage medium for implementing the method of any one of the first or second embodiments. The embodiments of the computer-readable storage medium provided in this application are described in a relatively simple manner; relevant parts can be found in the corresponding descriptions of the above method embodiments. The embodiments described below are merely illustrative.
[0144] The computer-readable storage medium provided in this embodiment stores computer instructions, which, when executed by a processor, implement the steps described in the first or second embodiment.
[0145] In a typical configuration, a computing device includes one or more processors (CPU), input / output interfaces, network interfaces, and memory.
[0146] Memory may include non-persistent storage in computer-readable media, such as random access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash RAM. Memory is an example of computer-readable media.
[0147] 1. Computer-readable media includes both permanent and non-permanent, removable and non-removable media, which can store information by any method or technology. Information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase-change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, CD-ROM, digital versatile optical disc (DVD) or other optical storage, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other non-transferable medium that can be used to store information accessible by a computing device. As defined in this application, computer-readable media does not include non-transitory computer-readable media, such as modulated data signals and carrier waves.
[0148] 2. Those skilled in the art will understand that embodiments of this application can provide methods, systems, or computer program products. Therefore, embodiments of this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, embodiments of this application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0149] 3. This application embodiment may involve the use of user data. In practical applications, user-specific personal data may be used within the scope permitted by applicable laws and regulations of the country in which the application is located (e.g., with the user's explicit consent and effective notification to the user, etc.). Furthermore, the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties. The collection, use and processing of related data must comply with relevant laws, regulations and standards, and corresponding operation entry points are provided for users to choose to authorize or refuse.
[0150] Although this application discloses preferred embodiments as described above, it is not intended to limit this application. Any person skilled in the art can make possible changes and modifications without departing from the spirit and scope of this application. Therefore, the scope of protection of this application should be determined by the scope defined in the claims of this application.
Claims
1. A method for determining image matting parameters, characterized in that, The method includes: Obtain the image to be cut out; Based on a pre-trained segmentation model, ground truth segmentation information is determined for each pixel in the image to be cut out. The ground truth segmentation information is used to indicate whether the corresponding pixel is the background or the foreground. Determine the matting parameters to be optimized, which are used to control the overall matting intensity of the image; The image to be cut out is rendered according to the cutout parameters to obtain the effect segmentation information corresponding to each pixel in the image to be cut out; With the optimization objective of reducing the loss between the effect segmentation information and the ground truth segmentation information, the matting parameters are adjusted to obtain the target matting parameters.
2. The method for determining image matting parameters according to claim 1, characterized in that, The optimization objective is to reduce the loss between the effect segmentation information and the ground truth segmentation information. The matting parameters are adjusted to obtain the target matting parameters, including: With the optimization goal of reducing the pixel-level difference between the effect segmentation information and the ground truth segmentation information, the matting parameters are adjusted in the first stage to obtain the first stage parameters; With the optimization objective of reducing the structural difference between the effect segmentation information and the ground truth segmentation information, the parameters of the first stage are adjusted in the second stage to obtain the target matting parameters.
3. The method for determining image matting parameters according to claim 2, characterized in that, The optimization objective is to reduce the pixel-level difference between the effect segmentation information and the ground truth segmentation information. The first-stage parameter adjustment of the matting parameters yields the following parameters: Based on the segmentation region corresponding to each pixel in the true value segmentation information, the loss weight corresponding to each pixel is determined, wherein the loss weight corresponding to the background segmentation region is greater than the loss weight corresponding to the foreground segmentation region. The pixel-level difference between the effect segmentation information and the ground value segmentation information is weighted according to the loss weight corresponding to each pixel to obtain the pixel-level weighted error. With minimizing the pixel-level weighted error as the optimization objective, the matting parameters are adjusted in the first stage to obtain the first-stage parameters.
4. The method for determining the matting parameters according to claim 3, characterized in that, The step of determining the loss weight corresponding to each pixel based on the segmentation region corresponding to each pixel in the truth segmentation information includes: Get the preset cutout area; The loss weight of each pixel in the cutout area is set to 0; Based on the segmentation region corresponding to each pixel in the true value segmentation information, the loss weight corresponding to each pixel in the region outside the cutout region in the image to be cut out is determined.
5. The method for determining image matting parameters according to claim 2, characterized in that, The second-stage parameter adjustment of the first-stage parameters, with the optimization objective of reducing the structural difference between the effect segmentation information and the ground truth segmentation information, includes: When there is a first connected component in the background region of the truth segmentation information that is determined to be the foreground by the effect segmentation information, the parameters of the first stage are adjusted in the second stage in the direction of increasing matting intensity to reduce the number of the first connected components. When there is a second connected component in the foreground region of the truth segmentation information that is determined to be the background by the effect segmentation information, and there is no first connected component in the background region of the truth segmentation information, the parameters of the first stage are adjusted in the second stage in the direction of decreasing matting intensity to reduce the number of the second connected components.
6. The method for determining image matting parameters according to claim 5, characterized in that, The second-stage parameter adjustment of the first-stage parameters in the direction of increasing matting intensity includes: With the condition that the number of third connected regions does not exceed a preset threshold, the parameters of the first stage are adjusted in the direction of increasing image matting intensity in the second stage. The third connected region is a connected region that is determined as background by the effect segmentation information within the foreground region of the ground truth segmentation information and intersects with the boundary of the foreground region of the ground truth segmentation information.
7. The method for determining image matting parameters according to claim 5, characterized in that, The second stage parameter adjustment of the first stage parameters in the direction of decreasing matting intensity includes: With the constraint that the image structure similarity between the effect segmentation information and the ground truth segmentation information is not lower than a preset similarity threshold, the parameters of the first stage are adjusted in the second stage in the direction of decreasing the matting intensity.
8. The method for determining image matting parameters according to claim 7, characterized in that, The image structural similarity is determined in the following way: Determine the standard structural similarity between the effect segmentation information and the truth segmentation information; The edge contours of the effect segmentation information and the ground truth segmentation information are extracted respectively to obtain a first edge map and a second edge map; Determine the edge structure similarity between the first edge map and the second edge map; Based on the standard structural similarity and the edge structural similarity, the image structural similarity between the effect segmentation information and the ground truth segmentation information is determined.
9. The method for determining image matting parameters according to claim 8, characterized in that, The step of determining the image structural similarity between the effect segmentation information and the ground truth segmentation information based on the standard structural similarity and the edge structural similarity includes: The background coverage integrity value is determined based on the number of first pixels in the background region of the true value segmentation information that are identified as the foreground region by the effect segmentation information, wherein the background coverage integrity value is negatively correlated with the number of first pixels; Using the background coverage integrity value as a weight, the standard structural similarity and the edge structural similarity are weighted to obtain the image structural similarity between the effect segmentation information and the ground truth segmentation information.
10. The method for determining matting parameters according to any one of claims 1 to 9, characterized in that, The segmentation model was trained in the following way: Acquire training samples, which include sample images and corresponding sample ground truth segmentation information, wherein the sample images are HSV color space images; At least one of the hue channel, saturation channel, and brightness channel of the sample image is randomly perturbed to generate a sample-enhanced image; The segmentation model is trained based on the sample image, the sample augmented image, and the corresponding sample ground truth segmentation information.
11. The method for determining matting parameters according to any one of claims 1 to 9, characterized in that, The image to be cut out is a video frame in the video to be cut out, and the target cutout parameters are the same for each video frame.
12. A method for image matting, characterized in that, The method includes: Obtain the video to be cut out, and determine each video frame in the video to be cut out as the image to be cut out; The method for determining matting parameters as described in any one of claims 1 to 11 is used to determine the target matting parameters corresponding to each video frame of the video to be matted. The video frames are rendered by matting according to the target matting parameters to generate the matted result image.
13. An electronic device, characterized in that, include: Processor, memory, and computer program instructions stored in said memory and executable on the processor; When the processor executes the computer program instructions, it implements the method as described in any one of claims 1-12.
14. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer-executable instructions, which, when executed by a processor, are used to implement the method as described in any one of claims 1-12.