Image matting method and device based on alpha mask refinement

By using the cutout model and the detail difference feature extractor on high-resolution images to generate high-resolution alpha masks and optimize the cutout model, the problem of high-resolution image cutout efficiency and low quality in the existing technology is solved, and efficient and accurate image cutouts are achieved.

CN120013981APending Publication Date: 2025-05-16GUIZHOU MINZU UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510054540.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-14
Publication Date
2025-05-16

AI Technical Summary

Technical Problem

The prior art is difficult to obtain high-quality alpha masks on high-resolution images within a reasonable time, resulting in reduced image cutout efficiency and quality.

Method used

By inputting high-resolution images as training samples into the preset cutout model, a low-resolution alpha mask is obtained, and the detailed difference feature extractor is used to obtain the detailed difference characteristics between the images, and the fusion process is performed to generate the target high-resolution alpha mask, while optimizing the cutout model with the loss function.

Benefits of technology

The clarity and accuracy of the target high-resolution alpha mask is significantly improved, avoiding the high computing cost brought by high-resolution pinching, and achieving dual optimization of high-resolution image pinching efficiency and quality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120013981A_ABST
    Figure CN120013981A_ABST
Patent Text Reader

Abstract

The invention discloses an image matting method and device based on alpha mask refinement, and the method comprises the steps: obtaining a plurality of training samples of a high-resolution image, inputting the training samples into a preset matting model, and obtaining a low-resolution alpha mask corresponding to the training samples; using a detail difference feature extractor to obtain detail difference features between the training samples and the corresponding low-resolution images; performing fusion processing on the detail difference features and the low-resolution alpha mask to obtain a target high-resolution alpha mask; and obtaining error information between a real alpha mask corresponding to the training sample and the target high-resolution alpha mask, and performing optimization processing on the preset matting model by using a preset loss function based on the error information to obtain a target matting model. According to the invention, a complex high-resolution image matting problem is converted into a low-resolution image matting problem and a high-resolution alpha mask refining problem, and double optimization of high-resolution image matting efficiency and quality is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of image processing technology, and in particular to an image cutout method and device based on alpha mask refinement. Background Art

[0002] High-resolution natural image matting is the process of precisely extracting the foreground from the background of a high-resolution image by accurately determining its opacity. It is essential in several key applications, such as image editing, filmmaking, and remote sensing. At higher resolutions, even small errors become very obvious, and this increased detail poses a challenge to matting algorithms. Therefore, as the number of pixels increases, accurately determining the alpha value of each pixel in a reasonable time becomes critical to maintain the efficiency and quality of image matting. However, due to the complexity brought about by the surge in resolution, existing image matting methods are unable to obtain high-quality alpha masks on high-resolution images in a reasonable time.

[0003] Therefore, how to improve the performance of high-resolution image matting methods is an urgent problem to be solved. Summary of the invention

[0004] In order to solve the above technical problems, an embodiment of the present application provides an image cutout method and device based on alpha mask refinement.

[0005] In a first aspect, in order to solve the above technical problems, the present application provides an image cutout method based on alpha mask refinement, comprising:

[0006] Acquire a plurality of training samples, wherein the training samples are high-resolution images;

[0007] Input the training sample into a preset cutout model to obtain a low-resolution alpha mask corresponding to the training sample;

[0008] Using a detail difference feature extractor, obtaining detail difference features between the training sample and the corresponding low-resolution image;

[0009] The detail difference feature and the low-resolution alpha mask are fused to obtain a target high-resolution alpha mask;

[0010] The error information between the real alpha mask corresponding to the training sample and the target high-resolution alpha mask is obtained, and the preset cutout model is optimized based on the error information using a preset loss function to obtain a target cutout model, which is applied to the cutout of a large number of high-resolution images.

[0011] The beneficial effects are:

[0012] In the technical solution provided in the embodiment of the present application, a high-resolution image is input as a training sample into a preset cutout model for model training, so as to obtain a target cutout model applied to the cutout of a large number of high-resolution images. First, a low-resolution alpha mask corresponding to the training sample is obtained, which can reduce the amount of calculation by reducing the image resolution, and at the same time generate a preliminary low-resolution alpha mask, and simplify the image cutout process without affecting the quality of the final alpha mask. Secondly, a detail difference feature extractor is used to obtain the detail difference features between the training sample and the corresponding low-resolution image. After that, the detail difference features and the low-resolution alpha mask are fused to obtain a target high-resolution alpha mask, that is, the low-resolution scene is refined by using the complex detail information of the high-resolution image, and the clarity and accuracy of the target high-resolution alpha mask are significantly improved by fusing and integrating these details. Finally, the error information between the real alpha mask corresponding to the training sample and the target high-resolution alpha mask is obtained, and the preset loss function is used to optimize the preset cutout model based on the error information to obtain the target cutout model. In this way, when the target cutout model is used to cut out a large number of high-resolution images, not only the quality of high-resolution alpha matting is guaranteed, but also the high computational cost brought by high-resolution matting is avoided, thereby achieving dual optimization of high-resolution image matting efficiency and quality.

[0013] In a second aspect, the present invention provides an image matting device based on alpha mask refinement, comprising an acquisition unit, a feature unit, a processing unit and a loss unit;

[0014] An acquisition unit, used for acquiring a plurality of training samples, wherein the training samples are high-resolution images;

[0015] A cutout unit, used for inputting the training sample into a preset cutout model to obtain a low-resolution alpha mask corresponding to the training sample;

[0016] A feature unit, configured to obtain detail difference features between the training sample and the corresponding low-resolution image using a detail difference feature extractor;

[0017] A processing unit, configured to fuse the detail difference feature with the low-resolution alpha mask to obtain a target high-resolution alpha mask;

[0018] The loss unit is used to obtain the error information between the real alpha mask corresponding to the training sample and the target high-resolution alpha mask, and optimize the preset cutout model based on the error information using a preset loss function to obtain a target cutout model, which is applied to the cutout of a large number of high-resolution images.

[0019] It should be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the present application. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] The drawings herein are incorporated into the specification and constitute a part of the specification, showing embodiments consistent with the present application, and together with the specification, are used to explain the principles of the present application. Obviously, the drawings described below are only some embodiments of the present application, and for those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative work. In the drawings:

[0021] Figure 1 is a flowchart of an image cutout method based on alpha mask refinement shown in an exemplary embodiment of the present application;

[0022] Figure 2 is a schematic diagram of a detail difference feature extractor in an exemplary embodiment of the present application;

[0023] Figure 3 It is a schematic diagram of a process of obtaining a target high-resolution alpha mask in an exemplary embodiment provided by the present application;

[0024] Figure 4 is a schematic diagram of a process of obtaining fourth error information matching a mask detail resolution difference loss function in an exemplary embodiment of the present application;

[0025] Figure 5 is a schematic diagram of image resolution distribution of an experimental data set in an exemplary embodiment of the present application;

[0026] Figure 6 It is a schematic diagram of visual comparison between the image matting method based on alpha mask refinement provided by the present application and other methods on the high-resolution image part in the Alphamatting dataset;

[0027] Figure 7 This is a schematic diagram of visual comparison between the image matting method based on alpha mask refinement provided by this application and other methods on the Transparent-460 dataset;

[0028] Figure 8 is a block diagram of an image cutout device based on alpha mask refinement shown in an exemplary embodiment of the present application;

[0029] Fig. 9 It is a structural diagram of a computer system suitable for implementing an electronic device of an embodiment of the present application. DETAILED DESCRIPTION

[0030] Here, exemplary embodiments will be described in detail, examples of which are shown in the accompanying drawings. When the following description refers to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The implementations described in the following exemplary embodiments do not represent all implementations consistent with the present application. Instead, they are only examples of devices and methods consistent with some aspects of the present application as detailed in the attached claims.

[0031] The block diagrams shown in the accompanying drawings are merely functional entities and do not necessarily correspond to physically independent entities. That is, these functional entities may be implemented in software form, or in one or more hardware modules or integrated circuits, or in different networks and / or processor devices and / or microcontroller devices.

[0032] The flowcharts shown in the accompanying drawings are only exemplary and do not necessarily include all the contents and operations / steps, nor must they be executed in the order described. For example, some operations / steps can be decomposed, and some operations / steps can be combined or partially combined, so the actual execution order may change according to actual conditions.

[0033] The term "multiple" as used in this application refers to two or more than two. "And / or" describes the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can mean: A exists alone, A and B exist at the same time, and B exists alone. The character " / " generally indicates that the related objects are in an "or" relationship.

[0034] In the related art, high-resolution natural image matting is the process of accurately extracting the foreground from the background of a high-resolution image by accurately determining the opacity of the foreground, which is essential in several key applications, such as image editing, filmmaking, and remote sensing. At higher resolutions, even small errors can become very obvious, and this increased detail poses a challenge to matting algorithms.

[0035] Image matting can be divided into traditional methods and deep learning-based methods. Traditional methods include propagation-based methods and optimization-based methods. Propagation-based methods usually involve pixel-by-pixel analysis to determine the degree of similarity. They are impractical for high-resolution images because the huge number of pixels leads to a large number of comparisons. Although optimization-based methods regard matting as a pixel-pair optimization problem, they also face significant challenges when dealing with high-resolution images due to the exponential growth of the complexity and size of the search space with the resolution. As one of the traditional methods, the propagation-based method propagates the alpha value from the known region to the unknown region by measuring the similarity between the unknown pixel and the known foreground and background pixels. On the other hand, the optimization-based method models the image matting problem as a pixel-pair optimization problem and estimates the alpha value by solving the optimization problem for each unknown pixel.

[0036] However, evaluating and selecting the best foreground and background color sample pairs requires a lot of computation, resulting in high computational complexity. Even with strategies to alleviate the pressure on computing resources, traditional image matting methods are still unable to process high-resolution images. These methods usually rely on complex similarity metrics and precise operations at the pixel level, which are not flexible and efficient enough, and therefore cannot be applied to large amounts of high-resolution images. In addition, in areas with rich details, they may have difficulty accurately distinguishing between foreground and background, resulting in unsatisfactory matting results. These shortcomings are particularly prominent in the field of modern image processing, especially in scenarios where a large number of high-resolution images need to be processed quickly and accurately.

[0037] In recent years, deep learning-based methods have made great progress in the field of image matting. Deep learning technology has achieved accurate alpha estimation on low-resolution images through advanced network architecture and attention mechanism, making significant contributions to the field of image matting. However, there are still technical challenges in processing high-resolution images. The detail refinement of high-resolution images requires not only the algorithm to accurately estimate the alpha value, but also to accurately distinguish the foreground and background in complex scenes while retaining the subtle features of the image. Therefore, the main challenge facing deep learning-based matting technology is to complete fine detail processing under reasonable computing resources. If too much emphasis is placed on resource optimization, key details of the image may be sacrificed, affecting the final matting quality. On the contrary, if extreme retention of details is pursued, the increased computational burden will affect processing efficiency.

[0038] In order to solve the above problems, the embodiments of the present application propose an image cutout method and device based on alpha mask refinement, an electronic device, and a computer-readable storage medium, which mainly relate to the image cutout technology based on alpha mask refinement included in the image processing technology. These embodiments will be described in detail below.

[0039] First see Figure 1 , Figure 1 This is a flowchart of an image cutout method based on alpha mask refinement shown in an exemplary embodiment of the present application. The method can be specifically executed by a server, which can be an independent server or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content delivery networks (CDNs), and big data and artificial intelligence platforms, etc., and is not limited here.

[0040] like Figure 1 As shown, in an exemplary embodiment, the image matting method based on alpha mask refinement may include steps S101 to S105, which are described in detail as follows:

[0041] Step S101, obtaining a plurality of training samples, where the training samples are high-resolution images.

[0042] Step S102: input the training sample into a preset cutout model to obtain a low-resolution alpha mask corresponding to the training sample.

[0043] Step S103: using a detail difference feature extractor to obtain detail difference features between the training sample and the corresponding low-resolution image.

[0044] Step S104: fusing the detail difference feature with the low-resolution alpha mask to obtain a target high-resolution alpha mask.

[0045] Step S105, obtaining the error information between the real alpha mask corresponding to the training sample and the target high-resolution alpha mask, and optimizing the preset cutout model based on the error information using a preset loss function to obtain a target cutout model, which is applied to the cutout of a large number of high-resolution images.

[0046] In this embodiment, high-resolution images are input as training samples into a preset cutout model for model training, so as to obtain a target cutout model applied to the cutout of a large number of high-resolution images.

[0047] Preferably, first obtain the low-resolution alpha mask corresponding to the training sample, so that the amount of calculation can be reduced by reducing the image resolution, and a preliminary low-resolution alpha mask is generated at the same time, and the image matting process is simplified without affecting the quality of the final alpha mask. Secondly, use a detail difference feature extractor to obtain the detail difference features between the training sample and the corresponding low-resolution image. After that, the detail difference features and the low-resolution alpha mask are fused to obtain a target high-resolution alpha mask, that is, the low-resolution scene is refined using the complex detail information of the high-resolution image, and the clarity and accuracy of the target high-resolution alpha mask are significantly improved by fusing and integrating these details. Finally, obtain the error information between the real alpha mask corresponding to the training sample and the target high-resolution alpha mask, and use a preset loss function to optimize the preset matting model based on the error information to obtain a target matting model, which is applied to the matting of a large number of high-resolution images.

[0048] Therefore, the present application transforms the complex high-resolution matting problem into a low-resolution matting problem and a high-resolution alpha mask refinement problem. In the embodiments provided in the present application, the framework for applying the preset matting model and the target matting model can be referred to as HRIMF-AMR (high-resolution image matting framework based on alpha matte refinementfrom low-resolution To resolution). HRIMF-AMR uses a detail difference feature extractor (DDFE, Detail Differenc Feature Extractor) module to extract fine details unique to high-resolution images, and then uses these details to refine the corresponding low-resolution alpha mask. In addition, a preset loss function is also used to improve the extraction of resolution-specific details to ensure that the detail differences between high-resolution and low-resolution images are effectively captured in the target high-resolution alpha mask. In this way, not only the quality of high-resolution alpha matting is guaranteed, but also the high computational cost caused by high-resolution matting is avoided, thereby achieving dual optimization of high-resolution image matting efficiency and quality.

[0049] In an exemplary embodiment provided by the present application, a low-resolution alpha mask corresponding to a training sample is obtained by using a preset cutout network, and the specific steps may include:

[0050] Input the training sample into the preset cutout model to obtain a high-resolution three-dimensional image corresponding to the training sample, wherein the high-resolution three-dimensional image includes a foreground, a background, and an area to be predicted;

[0051] The training samples and the high-resolution three-dimensional map are input into the preset cutout network to obtain the low-resolution alpha mask corresponding to the training samples.

[0052] The preset matting network provided in the present application may be Matteformer, which is a competitive matting method that helps to accurately extract low-resolution matting from low-resolution images. Preferably, Matteformer relies on trimap (static image matting algorithm) to roughly divide the image to distinguish between foreground, background and unknown areas to be predicted.

[0053] In this embodiment, after the training sample is input into the preset matting model, the trimap is used to obtain the high-resolution three-dimensional image corresponding to the training sample, and the high-resolution three-dimensional image includes the foreground, the background and the area to be predicted. The training sample as a high-resolution image and the high-resolution three-dimensional image are connected through a channel and input into the preset matting network Matteformer Network to obtain the low-resolution alpha mask corresponding to the training sample.

[0054] In this way, through the above-mentioned embodiments, the present application obtains a low-resolution alpha mask corresponding to the training sample based on the training sample of the high-resolution image and the corresponding high-resolution triplicate map in a preset cutout network, thereby simplifying the image cutout process without affecting the quality of the final alpha mask and further improving the image processing efficiency.

[0055] In an exemplary embodiment provided by the present application, the specific steps of obtaining detail difference features between a training sample and a corresponding low-resolution image may include:

[0056] Using the detail difference feature extractor, the training samples are downsampled to obtain low-resolution images;

[0057] Extracting features from the training sample and the low-resolution image respectively to obtain a first detail feature and a second detail feature;

[0058] Distinguishing the first detail feature and the second detail feature to obtain a distinguishing feature;

[0059] The distinguishing features are processed by dimensionality reduction and normalization in turn to obtain the detailed difference features between the training samples and the low-resolution images.

[0060] Since high-resolution images contain a large number of pixels, they capture more details than low-resolution images, allowing for finer local variations and texture information. This enhanced detail is crucial for high-resolution image matting because it provides rich visual cues that can significantly improve the accuracy of the matting process.

[0061] Therefore, in this embodiment, in order to extract the detail difference between the high-resolution and low-resolution images, the dimension of the high-resolution image is firstly reduced by downsampling to obtain a low-resolution image. Considering that the high-resolution natural image matting needs to minimize the computational complexity, the nearest neighbor method is preferably used for downsampling. Then, the detail difference feature extractor is used to extract features of the training sample and the low-resolution image respectively to obtain the first detail feature and the second detail feature, and the first detail feature and the second detail feature are distinguished to obtain the distinguishing feature. Finally, the distinguishing feature is sequentially subjected to dimensionality reduction and normalization processing to obtain the detail difference feature between the training sample and the low-resolution image.

[0062] See also Figure 2 , Figure 2 FIG. 1 is a schematic diagram of a detail difference feature extractor in an exemplary embodiment of the present application. Figure 2 As shown, the detail difference feature extractor (DDFE) consists of a series of convolutional layers for comparing details from a high-resolution image with details from a low-resolution image, with the aim of capturing details in the high-resolution image.

[0063] From the structure of DDFE, we can see that calculating detail difference features includes 3 steps. First, a simple convolutional neural network layer is applied on the high-resolution image, aiming to capture its fine first detail features with a simple structure. Next, the high-resolution image is downsampled to generate a low-resolution image and extract its second detail features. Finally, by comparing the features extracted from the high-resolution image and the low-resolution image, unique details that only appear in the high-resolution image are extracted. These steps are completed using the following formula:

[0064] HDF=Conv(I h )

[0065] LDF=Conv T (I l )

[0066] DDF=Conv(HDF-LDF)

[0067] DDF offset =Tanh(Conv 1×1 (DDF)

[0068] In the formula, I h and I l Represent high-resolution images and low-resolution images respectively, HDF and LDF represent the detail features of high-resolution images and low-resolution images respectively, and DDF offset Represents the detail difference features between high-resolution images and low-resolution images. Conv and Conv 1×1 It is a simple convolutional layer, each of which plays a specific role. Use Conv for feature extraction and Conv 1×1 To reduce the dimension, use Conv in the Conv Transpose Layer T In order to limit the feature value to the normalized range, the calculation result is normalized by applying the Tanh function. This step ensures that the range of DDF is between -1 and 1, thereby generating an offset for alpha mask refinement, that is, the detail difference feature between the training sample and the low-resolution image, which is used to correct the low-resolution alpha mask by fusion.

[0069] In this way, through the above-mentioned embodiments, the detail difference feature extractor DDFE of the present application provides foreground details for high-resolution mask refinement by capturing subtle detail differences between high-resolution and low-resolution images to integrate these features with the low-resolution alpha mask in the cutout network to reconstruct a more accurate high-resolution mask as the target high-resolution alpha mask.

[0070] In an exemplary embodiment provided by the present application, the step of obtaining a target high-resolution alpha mask by fusing the detail difference feature and the low-resolution alpha mask may specifically include:

[0071] Perform upsampling on the low-resolution alpha mask to obtain a processed low-resolution alpha mask;

[0072] The processed low-resolution alpha mask and detail difference features are fused to obtain the target high-resolution alpha mask.

[0073] In this embodiment, the final alpha mask is predicted by fusion of the upsampled low-resolution alpha mask and the mask detail difference features, so as to refine the low-resolution scene using the complex detail information of the high-resolution image. By fusing and integrating these details, the clarity and accuracy of the target high-resolution alpha mask can be improved.

[0074] See also Figure 3 , Figure 3 FIG. 1 is a flow chart of obtaining a target high-resolution alpha mask in an exemplary embodiment provided by the present application. Figure 3 As shown, the high-resolution triplicated image corresponding to the training sample is obtained, and the training sample as the high-resolution image and the high-resolution triplicated image are connected through a channel. After downsampling, they are input into the preset matting network Matteformer Network to obtain the low-resolution alpha mask corresponding to the training sample. In addition, the detail difference feature extractor is used to extract features to obtain the detail difference features between the training sample and the low-resolution image. After that, the processed low-resolution alpha mask obtained by upsampling is subjected to feature fusion processing with the detail difference features to obtain the target high-resolution alpha mask.

[0075] In the field of image processing, especially in the context of high-resolution image matting, the transition from low resolution to high resolution is not just a matter of enlargement, but involves capturing complex details that are often lost in the process. Therefore, this application uses a variety of loss functions in the stage of training a preset matting model, including regression loss function, synthesis loss function, Laplacian loss function and Matte Detail Resolution Difference (MDRD) loss function.

[0076] In an exemplary embodiment provided by the present application, the specific steps of optimizing the preset cutout model based on the error information using the preset loss function to obtain the target cutout model may include:

[0077] Obtaining first error information between the real alpha mask corresponding to the training sample and the target high-resolution alpha mask that matches the regression loss function;

[0078] Obtaining second error information between the true alpha mask and the target high-resolution alpha mask that matches the synthesis loss function;

[0079] Obtain the third error information between the true alpha mask and the target high-resolution alpha mask that matches the Laplace loss function;

[0080] Obtain fourth error information between the true alpha mask and the target high-resolution alpha mask that matches the mask detail resolution difference loss function;

[0081] The preset cutout model is optimized based on the corresponding first error information, second error information, third error information and fourth error information by using the regression loss function, the synthesis loss function, the Laplace loss function and the difference in mask detail resolution to obtain the target cutout model.

[0082] In another exemplary embodiment, the specific step of obtaining first error information between a real alpha mask corresponding to the training sample and a target high-resolution alpha mask that matches the regression loss function may include:

[0083] Obtaining first mask information at the to-be-predicted area in the real alpha mask corresponding to the training sample, and second mask information at the to-be-predicted area in the target high-resolution alpha mask;

[0084] The mean absolute error between the first mask information and the second mask information is obtained as the first error information matched with the regression loss function.

[0085] In this embodiment, the calculation formula for obtaining the mean absolute error between the first mask information and the second mask information is:

[0086]

[0087] in, represents the mean absolute error, It is the area to be predicted marked by trimap. and α i They represent the target high-resolution alpha mask and the real alpha mask at position i respectively.

[0088] In another exemplary embodiment, the specific step of obtaining second error information between the real alpha mask and the target high-resolution alpha mask that matches the synthesis loss function may include:

[0089] Obtain a first image synthesized by a real alpha mask on the foreground and background, and a second image synthesized by a target high-resolution alpha mask on the foreground and background;

[0090] An absolute difference between the first image and the second image is obtained as second error information matched with the synthetic loss function.

[0091] In the depth image matting, the synthesis loss is defined as the absolute difference between the RGB image colors synthesized by the predicted high-resolution alpha mask and the real alpha mask on the ground-truth foreground and ground-truth background, respectively. In this embodiment, a first image synthesized by the real alpha mask on the foreground and background, and a second image synthesized by the target high-resolution alpha mask on the foreground and background are obtained, and the absolute difference between the first image and the second image is calculated. The calculation formula of the absolute difference is:

[0092]

[0093] in, represents the absolute difference, c pre The second image corresponding to the target high-resolution alpha mask, c gt Represents the first image corresponding to the true alpha mask, ∈ represents a small number, usually a constant to avoid synthesis loss of zero.

[0094] In another exemplary embodiment, the specific step of obtaining third error information between the real alpha mask and the target high-resolution alpha mask that matches the Laplace loss function may include:

[0095] The Laplace difference between the true alpha mask and the target high-resolution alpha mask is obtained as the third error information matched with the Laplace loss function.

[0096] Laplacian Loss is mainly used to minimize the difference in color information. Laplacian loss plays a role in video interpolation and cutout tasks. It captures local and global differences in images by constructing Gaussian pyramids and Laplacian pyramids.

[0097] In this example, the Laplacian loss is defined as the difference between the target high-resolution alpha mask and the true alpha mask as follows:

[0098]

[0099] in, Expressed as Laplace difference, L i represents the i-th level of the Laplacian pyramid of the alpha mask, and α represent the target high-resolution alpha mask and the true alpha mask respectively.

[0100] In another exemplary embodiment, the specific step of obtaining fourth error information between the real alpha mask and the target high-resolution alpha mask that matches the mask detail resolution difference loss function may include:

[0101] Get the difference in mask details between the real alpha mask and the low-resolution alpha mask;

[0102] Fourth error information matching the mask detail resolution difference loss function is obtained based on the mask detail difference and the detail difference feature output by the detail difference feature extractor.

[0103] See also Figure 4 , Figure 4 FIG. 1 is a flow chart of obtaining fourth error information matching the mask detail resolution difference loss function in an exemplary embodiment of the present application. Figure 4As shown, the low-resolution alpha mask output by the preset matting network MatteformerNetwork is obtained, and the low-resolution alpha mask is upsampled and then distinguished from the real alpha mask to extract the mask detail difference between the real alpha mask and the low-resolution alpha mask. In addition, the regional weight of the area to be predicted in the high-resolution tri-image corresponding to the training sample is first determined, and the feature information of the area to be predicted is determined from the detail difference features output by the detail difference feature extractor based on the regional weight, which can be achieved by means of masking. Then, the feature information and the mask detail difference are calculated to obtain the Euclidean norm of the corresponding value as the fourth error information, and the calculation formula is as follows:

[0104]

[0105] in, Indicates the fourth error information, DDF offset and Diff α They are denoted as detail difference feature and mask detail difference respectively, and ||*||2 is defined as the Euclidean norm.

[0106] The MDRD (Matte Detail Resolution Difference) loss function is a complement to the total loss, where MDRD supervises DDFE to focus on extracting mask details that only exist in high-resolution images, regression loss supervises the overall quality of high-resolution masks, component loss supervises color accuracy in image synthesis, and Laplacian loss supervises the structural content of different alpha mask layers. Combining multiple losses can achieve the effect of accurately extracting high-resolution alpha mask details.

[0107] Integrating MDRD into total loss The calculation is expressed as:

[0108]

[0109] in, represents the total loss, represents the mean absolute error, represents the absolute difference, Expressed as Laplace difference, represents the fourth error information, and λ1, λ2, λ3, and λ4 respectively represent weights corresponding to the respective error information.

[0110] In this way, through the above embodiments, the present application uses preset loss functions such as regression loss function, synthesis loss function, Laplace loss function and mask detail resolution difference loss function to introduce additional constraints for feature extraction to ensure that the extracted features can represent the detail differences.

[0111] In another exemplary embodiment provided by the present application, in order to reflect the practical utility of the image matting method based on alpha mask refinement of the present application, the image matting method based on alpha mask refinement integrated with the present application (HRIMF-AMR method) is compared quantitatively and qualitatively with the most advanced methods when applied to high-resolution datasets. The fused HRIMF-AMR method is applied to high-resolution images and compared with other SOTA methods.

[0112] The experimental data set of this embodiment uses the Alphamatting part and Transparent-460 data set with high-resolution images, and their image resolution is usually 2K or above. Figure 5 As shown, Figure 5 It is a schematic diagram of the image resolution distribution of the experimental data set in an exemplary embodiment of the present application.

[0113] (1)Transparent-460

[0114] The Transparent-460 dataset contains 460 well-annotated high-fidelity alpha mattes, of which 410 images are in the training subset and 50 images are in the test subset. The background is Adobe Composition-1K based on the Microsoft COCO and PASCAL VOC2012 datasets, which contains 41,000 training samples and 1,000 test samples. In addition, the images comprising the Transparent-460 test dataset have a larger resolution, with an average size of 3915×4059 pixels (ranging from 1661×1661 to 4480×6720 pixels).

[0115] (2) Alphamatting

[0116] The Alphamatting dataset consists of 8 test images and 27 training images, but both high-resolution and low-resolution versions are provided for each image. Since the Adobe Composition-1K dataset integrates images from Alphamatting-train, only 8 images from Alphamatting-test are used in the experiments. The average resolution of these 8 images is 3127×2364 pixels, ranging from 2689×2085 to 3908×2600 pixels. However, the average resolution of the corresponding low-resolution images is 800×607 (the minimum is 800×532 and the maximum is 800×671 pixels). All images in this dataset are natural (i.e., there are no synthetic images).

[0117] (3) Adobe Composition-1K

[0118] The Adobe Composition-1K dataset contains 43,100 images, each with an alpha mask attached, consisting of 431 different foreground elements fused with a corresponding number of background images. These backgrounds are randomly selected from the Microsoft COCO collection. In addition, the Adobe Composition-1K test dataset contains 1,000 images that are a mixture of 50 unique foreground images with backgrounds from the Pascal VOC 2012 dataset. In this test dataset, the average resolution of images ranges from 1120×502 to 1920×1920 pixels, with an average of 1655×1380 pixels.

[0119] After applying the image matting method based on alpha mask refinement and other methods of the present application to the above dataset, the performance of HRIMF-AMR is evaluated by five indicators, four of which measure the quality of predicted alphamattes and the remaining one reflects the computational complexity. Preferably, the evaluation indicators include the sum of absolute differences, mean square error, gradient error and connectivity error.

[0120] The Sum of Absolute Differences (SAD) metric measures the overall error by calculating the sum of the absolute differences of all pixels between the predicted alpha mask (i.e. the target high-resolution alpha mask) and the true alpha mask. It is the most intuitive error metric because it directly accumulates the prediction error of each pixel. The expression for calculating SAD is as follows:

[0121]

[0122] where α iis the predicted ith alpha mask pixel value, is the i-th pixel value of the true alpha mask, and Ω is the total number of pixels.

[0123] The mean squared error (MSE) metric emphasizes the impact of larger errors by squaring the prediction errors and then averaging them. This metric makes the contribution of a single large error to the overall error more significant, which helps to identify and improve significant defects in the algorithm. The calculation formula for MSE is:

[0124]

[0125] For clarity, this paper defines MSE as 10 -3 .

[0126] The gradient error (Grad) metric focuses on evaluating the edge quality of the predicted alpha scene by calculating the difference between the gradient of the predicted alpha scene and the true alpha scene. This metric is particularly important for matting tasks that require high edge details because it reflects the smoothness and accuracy of the edge. The calculation formula for the gradient error is:

[0127]

[0128] in and are the gradients of the predicted alpha mask and the true alpha mask, respectively, and q is a constant, usually 2.

[0129] The connectivity error (Conn) measures the difference between the connectivity of the foreground pixels in the predicted alpha mask and the connectivity in the true alpha mask. This index is crucial to ensure the integrity and connectivity of the foreground objects in the image matting. The calculation formula of the connectivity error is:

[0130]

[0131] in and denote the connectivity of pixel i in the predicted alpha mask and the true alpha mask, respectively, and Ω denotes the foreground area.

[0132] Latency is a key performance indicator for measuring system response time, reflecting the total time from issuing a request to receiving a response. Latency is one of the most intuitive indicators for measuring system performance, especially real-time system performance, because it directly accumulates the processing time of each request. For ease of analysis, this article converts it into average latency (AL, Average Logistics), calculated using the following expression:

[0133]

[0134] in is the time point when the i-th request is issued, is the time point when the ith request receives a response, and m is the total number of requests in the test.

[0135] For the detailed settings in the experimental process, in order to provide an alpha mask for comparison of image matting quality, HRIMF-AMR is integrated into Matteformer, and all the methods involved are implemented using PyTorch. It should be noted that the experiments discussed in this embodiment (unless otherwise stated) are conducted under the same resource configuration as the server cluster, that is, using 6 core CPUs and 1 GPU. In order to make a fair comparison with other state-of-the-art methods, the method integrated with HRIMF-AMR is also trained on the Adobe Composition-1K dataset. During the training phase, the input images are randomly cropped to 1024×1024. Then, these images are subjected to random affine transformation, cropping, and real-world enhancement according to the strategy described in Mask-Guided-Matting. Affine transformations include random rotation, scaling, cropping, and vertical and horizontal flipping. In order to speed up the training process and prevent overfitting problems, the pre-trained weights of Matteformer are used to train the entire model in an end-to-end manner. For loss optimization, we use an optimizer with β1=0.5 and β2=0.999 settings. Furthermore, the initial optimizer learning rate was set to 1e-3 and a warm-up strategy was used for 2500 iterations, after which the learning rate was adjusted according to the decay law of the learning rate. In the fine-tuning phase, the optimal HRIMF-AMR model was trained on NVIDIA A100 for 16 batches and 50,000 iterations. In the inference phase, high-resolution images and trimaps were input into the network to predict high-resolution alpha masks.

[0136] From this experiment, it can be seen that the quantitative results obtained by applying the image matting method based on alpha mask refinement of this application on Transparent-460 and the most advanced method when applied to high-resolution datasets are shown in the following Table 1. Table 1:

[0137]

[0138] *Graphics memory requirements exceed40GB.

[0139] Experimental results show that HRIMF-AMR is able to extract detail information of high-resolution images within a reasonable time, that is, the image matting method based on alpha mask refinement of the present application provides high-quality high-resolution alpha masks within a reasonable time.

[0140] See also Figure 6 and Figure 7 , Figure 6 is a schematic diagram of visual comparison between the image matting method based on alpha mask refinement provided by this application and other methods on the high-resolution image part in the Alphamatting dataset. Figure 7 : is a visual comparison diagram of the image matting method based on alpha mask refinement provided by this application and other methods on the Transparent-460 dataset. Figure 6 As shown, the image matting method based on alpha mask refinement provided by the present application is still very competitive in detail extraction of high-resolution natural images, and can extract fine details such as glass reflection and noise of transparent spheres from high-resolution images.

[0141] In order to verify the adaptability of the HRIMF-AMR framework proposed in this application, it is compared with DIM (Deep ImageMatting, an image matting technology based on deep learning), A 2 The classical methods such as U, MGM-trimap and Matteformer are integrated. These methods have been carefully integrated, trained and tested to ensure that the framework applied in this application works best in different cutout methods. The adaptive results of HRIMF-AMR are shown in Table 2 below. Table 2:

[0142]

[0143] As can be seen from Table 2, the overall performance of these classic methods after integration with HRIMF-AMR is significantly improved. Specifically, the average latency of the DIM method is reduced by about 39.3%, from 2.525 seconds to 1.535 seconds, while improving the accuracy, as SAD, MSE, Grad, and Conn are significantly reduced (21.4%, 36.3%, 18.1%, and 19.8%, respectively). HRIMF-AMR also solves the memory limitations encountered by some algorithms. 2The U method is a good example, which initially required more graphics memory than was available on the GPU device with HRIMF-AMR. By combining it with HRIMF-AMR, it was able to operate within these limitations. In addition, when merged into the HRIMF-AMR ensemble, the MGM-trimap method showed a 20.7% reduction in average latency, a 9.9% reduction in SAD, a 36.2% reduction in MSE, and a 15.7% reduction in Conn. The Matteformer method also showed significant efficiency improvements due to the adoption of HRIMF-AMR, with a 28.7% reduction in average latency and an increase in accuracy of 10.9% (SAD), 39.4% (MSE), and 0.03% (Conn). These results confirm that the HRIMF-AMR provided in this application can improve the accuracy of existing high-resolution image matting methods while reducing the computational resources required to process high-resolution images.

[0144] In addition, ablation experiments can be performed on the HRIMF-AMR framework of the image matting method based on alpha mask refinement provided by this application, by training on the Adobe Composition-1K dataset and testing on the Transparent-460 dataset. In this test, three schemes can be included: training and testing using interpolation-based upsampling and downsampling, using our DDFE, and using MDRD loss as an additional constraint. The test results are shown in Table 3 below. Table 3:

[0145]

[0146] The results shown in Table 3 show that using interpolation-based upsampling and downsampling for training and testing results in a considerable loss of detail information in high-resolution images, which is unacceptable for dense prediction tasks such as high-resolution image matting. Therefore, the reported results reflect the effectiveness of DDFE and MDRD with HRIMF-AMR. DDFE significantly reduces the SAD and MSE values, reflecting the improvement in accuracy in capturing details. The addition of the MDRD loss function further refines these metrics, indicating its subtle but positive impact on details. Since the gradient metric evaluates edge quality by comparing the gradient of the predicted alpha matte with the gradient of the true alpha matte, it initially increases with the addition of DDFE, indicating a temporary reduction in edge smoothness. Although the introduction of the MDRD loss improves the balance between detail preservation and edge smoothing, there is still room for further optimization to achieve the desired edge quality. For the Conn metric, which is crucial to ensure the integrity and connectivity of foreground objects, improvements over the DDFE module are shown, highlighting the role of MDRD in maintaining foreground connectivity. The MDRD loss further optimizes the Conn metric and strengthens the framework's ability to maintain the consistency of the alpha matte structure.

[0147] Figure 8 FIG. 8 is a block diagram of an image cutout device 800 based on alpha mask refinement, shown in an exemplary embodiment of the present application. Figure 8 As shown, the device comprises:

[0148] An acquisition unit 801 is used to acquire a plurality of training samples, where the training samples are high-resolution images;

[0149] A cutout unit 802 is used to input the training sample into a preset cutout model to obtain a low-resolution alpha mask corresponding to the training sample;

[0150] A feature unit 803 is used to obtain detail difference features between the training sample and the corresponding low-resolution image using a detail difference feature extractor;

[0151] A processing unit 804 is used to fuse the detail difference feature with the low-resolution alpha mask to obtain a target high-resolution alpha mask;

[0152] The loss unit 805 is used to obtain the error information between the real alpha mask corresponding to the training sample and the target high-resolution alpha mask, and optimize the preset cutout model based on the error information using a preset loss function to obtain a target cutout model, which is applied to the cutout of a large number of high-resolution images.

[0153] The device applies the image matting method based on alpha mask refinement provided by the present application, and inputs the high-resolution image transmitted by the acquisition unit 801 as a training sample into a preset matting model for model training, so as to obtain a target matting model applied to the matting of a large number of high-resolution images. Preferably, the low-resolution alpha mask corresponding to the training sample is obtained by the matting unit 802, so that the amount of calculation can be reduced by reducing the image resolution, and a preliminary low-resolution alpha mask is generated at the same time, and the image matting process is simplified without affecting the quality of the final alpha mask. Secondly, the feature unit 803 uses a detail difference feature extractor to obtain the detail difference feature between the training sample and the corresponding low-resolution image. Afterwards, the processing unit 804 fuses the detail difference feature and the low-resolution alpha mask to obtain a target high-resolution alpha mask, that is, the low-resolution scene is refined by using the complex detail information of the high-resolution image, and the clarity and accuracy of the target high-resolution alpha mask are significantly improved by fusing and integrating these details. Finally, the loss unit 805 obtains the error information between the real alpha mask corresponding to the training sample and the target high-resolution alpha mask, and uses the preset loss function to optimize the preset cutout model based on the error information to obtain the target cutout model. In this way, when using the target cutout model to cut out a large number of high-resolution images, not only the quality of high-resolution alpha matting is guaranteed, but also the high computational cost brought by high-resolution matting is avoided, thereby achieving dual optimization of high-resolution image matting efficiency and quality.

[0154] In another exemplary embodiment, the cutout unit 802 is further used to input the training sample into a preset cutout model to obtain a high-resolution tripartite map corresponding to the training sample, wherein the high-resolution tripartite map includes a foreground, a background, and an area to be predicted; and input the training sample and the high-resolution tripartite map into a preset cutout network to obtain a low-resolution alpha mask corresponding to the training sample.

[0155] In another exemplary embodiment, the feature unit 803 is also used to use a detail difference feature extractor to downsample the training sample to obtain a low-resolution image; perform feature extraction on the training sample and the low-resolution image respectively to obtain a first detail feature and a second detail feature; perform differentiation processing on the first detail feature and the second detail feature to obtain a distinguishing feature; perform dimensionality reduction processing and normalization processing on the distinguishing feature in turn to obtain a detail difference feature between the training sample and the low-resolution image.

[0156] In another exemplary embodiment, the processing unit 8041 is further used to upsample the low-resolution alpha mask to obtain a processed low-resolution alpha mask; and perform feature fusion processing on the processed low-resolution alpha mask and detail difference features to obtain a target high-resolution alpha mask.

[0157] In another exemplary embodiment, the preset loss function includes a regression loss function, a synthesis loss function, a Laplace loss function and a mask detail resolution difference loss function; the loss unit 805 is also used to obtain first error information between the real alpha mask and the target high-resolution alpha mask corresponding to the training sample that matches the regression loss function; obtain second error information between the real alpha mask and the target high-resolution alpha mask that matches the synthesis loss function; obtain third error information between the real alpha mask and the target high-resolution alpha mask that matches the Laplace loss function; obtain fourth error information between the real alpha mask and the target high-resolution alpha mask that matches the mask detail resolution difference loss function; and use the regression loss function, the synthesis loss function, the Laplace loss function and the mask detail resolution difference to optimize the preset cutout model based on the corresponding first error information, second error information, third error information and fourth error information to obtain the target cutout model.

[0158] In another exemplary embodiment, the loss unit 805 is also used to obtain first mask information at the to-be-predicted area in the real alpha mask corresponding to the training sample, and second mask information at the to-be-predicted area in the target high-resolution alpha mask; and obtain the mean absolute error between the first mask information and the second mask information as the first error information matching the regression loss function.

[0159] In another exemplary embodiment, the loss unit 805 is also used to obtain a first image synthesized by a real alpha mask on the foreground and background, and a second image synthesized by a target high-resolution alpha mask on the foreground and background; and obtain the absolute difference between the first image and the second image as second error information matching the synthesis loss function.

[0160] In another exemplary embodiment, the loss unit 805 is further configured to obtain a Laplace difference between the true alpha mask and the target high-resolution alpha mask as third error information matched with the Laplace loss function.

[0161] In another exemplary embodiment, the loss unit 805 is also used to obtain the mask detail difference between the real alpha mask and the low-resolution alpha mask; and obtain fourth error information matching the mask detail resolution difference loss function based on the mask detail difference and the detail difference feature output by the detail difference feature extractor.

[0162] It should be noted that the image cutout device based on alpha mask refinement provided in the above embodiment and the image cutout method based on alpha mask refinement provided in the above embodiment belong to the same concept, wherein the specific manner in which each module and unit performs the operation has been described in detail in the method embodiment and will not be repeated here. In practical applications, the image cutout device based on alpha mask refinement provided in the above embodiment can distribute the above functions to different functional modules as needed, that is, divide the internal structure of the device into different functional modules to complete all or part of the functions described above, and this is not limited here.

[0163] An embodiment of the present application also provides an electronic device, comprising: one or more processors; a storage device for storing one or more programs, and when the one or more programs are executed by one or more processors, the electronic device implements the image cutout method based on alpha mask refinement provided in the above-mentioned embodiments.

[0164] Fig. 9 The structure diagram of the computer system suitable for implementing the electronic device of the embodiment of the present application is shown. It should be noted that: Fig. 9 The computer system 900 of the electronic device shown is only an example and should not bring any limitation to the functions and scope of use of the embodiments of the present application.

[0165] like Fig. 9 As shown, the computer system 900 includes a central processing unit (CPU) 901, which can perform various appropriate actions and processes according to the program stored in the read-only memory (ROM) 902 or the program loaded from the storage part 908 to the random access memory (RAM) 903, such as executing the method in the above embodiment. In RAM 903, various programs and data required for system operation are also stored. CPU 901, ROM 902 and RAM 903 are connected to each other through a bus 904. Input / output (I / O) interface 905 is also connected to bus 904.

[0166] The following components are connected to the I / O interface 905: an input section 906 including a keyboard, a mouse, etc.; an output section 907 including a cathode ray tube (CRT), a liquid crystal display (LCD), etc., and a speaker, etc.; a storage section 908 including a hard disk, etc.; and a communication section 909 including a network interface card such as a LAN (Local Area Network) card, a modem, etc. The communication section 909 performs communication processing via a network such as the Internet. A drive 910 is also connected to the I / O interface 905 as needed. A removable medium 911, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on the drive 910 as needed so that a computer program read therefrom is installed into the storage section 908 as needed.

[0167] In particular, according to an embodiment of the present application, the process described above with reference to the flowchart can be implemented as a computer software program. For example, an embodiment of the present application includes a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program includes a computer program for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from a network through a communication section 909, and / or installed from a removable medium 911. When the computer program is executed by a central processing unit (CPU) 901, various functions defined in the system of the present application are executed.

[0168] It should be noted that the computer-readable medium shown in the embodiment of the present application can be a computer-readable signal medium or a computer-readable storage medium or any combination of the above two. The computer-readable storage medium can be, for example, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or device, or any combination of the above. More specific examples of computer-readable storage media can include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM), a flash memory, an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present application, a computer-readable signal medium may include a data signal propagated in a baseband or as part of a carrier wave, wherein a computer-readable computer program is carried. This propagated data signal can take a variety of forms, including but not limited to an electromagnetic signal, an optical signal, or any suitable combination of the above. A computer-readable signal medium may also be any computer-readable medium other than a computer-readable storage medium, which may send, propagate or transmit a program for use by or in conjunction with an instruction execution system, apparatus or device. A computer program contained on a computer-readable medium may be transmitted using any appropriate medium, including but not limited to: wireless, wired, etc., or any suitable combination of the above.

[0169] The flowcharts and block diagrams in the accompanying drawings illustrate the possible implementation architecture, functions and operations of the apparatus, methods and computer program products according to various embodiments of the present application. Among them, each box in the flowchart or block diagram can represent a module, a program segment, or a part of the code, and the above-mentioned module, program segment, or a part of the code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in an order different from that marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram or flowchart, and the combination of boxes in the block diagram or flowchart, can be implemented with a dedicated hardware-based system that performs a specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.

[0170] The units involved in the embodiments described in this application may be implemented by software or hardware, and the units described may also be set in a processor. The names of these units do not, in some cases, constitute limitations on the units themselves.

[0171] Another aspect of the present application further provides a computer-readable storage medium on which a computer program is stored, and when the computer program is executed by a processor, the image matting method based on alpha mask refinement as described above is implemented. The computer-readable storage medium may be included in the electronic device described in the above embodiment, or may exist independently without being assembled into the electronic device.

[0172] Another aspect of the present application also provides a computer program product or a computer program, which includes a computer instruction stored in a computer-readable storage medium. A processor of a computer device reads the computer instruction from the computer-readable storage medium, and the processor executes the computer instruction, so that the computer device executes the image matting method based on alpha mask refinement provided in each of the above embodiments.

[0173] The above are only preferred embodiments of the present application and are not intended to limit the present application. Any modifications, equivalent substitutions or improvements made within the spirit and principles of the present application should be included in the protection scope of the present application.

Claims

1. An image cutout method based on alpha mask refinement, characterized in that: The method comprises: Acquire a plurality of training samples, wherein the training samples are high-resolution images; Input the training sample into a preset cutout model to obtain a low-resolution alpha mask corresponding to the training sample; Using a detail difference feature extractor, obtaining detail difference features between the training sample and the corresponding low-resolution image; The detail difference feature and the low-resolution alpha mask are fused to obtain a target high-resolution alpha mask; The error information between the real alpha mask corresponding to the training sample and the target high-resolution alpha mask is obtained, and the preset cutout model is optimized based on the error information using a preset loss function to obtain a target cutout model, which is applied to the cutout of a large number of high-resolution images.

2. The method according to claim 1, characterized in that The step of inputting the training sample into a preset cutout model to obtain a low-resolution alpha mask corresponding to the training sample includes: Input the training sample into a preset cutout model to obtain a high-resolution three-part image corresponding to the training sample, wherein the high-resolution three-part image includes a foreground, a background, and a region to be predicted; The training sample and the high-resolution three-dimensional image are input into a preset cutout network to obtain a low-resolution alpha mask corresponding to the training sample.

3. The method according to claim 1 or 2, characterized in that: The step of using a detail difference feature extractor to obtain detail difference features between the training sample and the corresponding low-resolution image includes: Using a detail difference feature extractor, downsampling the training samples to obtain a low-resolution image; Extracting features from the training sample and the low-resolution image respectively to obtain a first detail feature and a second detail feature; Distinguishing the first detail feature and the second detail feature to obtain a distinguishing feature; The distinguishing features are sequentially subjected to dimensionality reduction processing and normalization processing to obtain detail difference features between the training sample and the low-resolution image.

4. The method according to claim 1, characterized in that: The step of fusing the detail difference feature with the low-resolution alpha mask to obtain a target high-resolution alpha mask comprises: Performing upsampling processing on the low-resolution alpha mask to obtain a processed low-resolution alpha mask; The processed low-resolution alpha mask and the detail difference feature are subjected to feature fusion processing to obtain a target high-resolution alpha mask.

5. The method according to claim 1, characterized in that The preset loss functions include a regression loss function, a synthesis loss function, a Laplace loss function, and a mask detail resolution difference loss function; The step of obtaining error information between the real alpha mask corresponding to the training sample and the target high-resolution alpha mask, and optimizing the preset cutout model based on the error information using a preset loss function to obtain a target cutout model includes: Acquire first error information between a real alpha mask corresponding to the training sample and the target high-resolution alpha mask that matches the regression loss function; Acquire second error information between the real alpha mask and the target high-resolution alpha mask that matches the synthesis loss function; Acquire third error information between the real alpha mask and the target high-resolution alpha mask that matches the Laplace loss function; Acquire fourth error information between the real alpha mask and the target high-resolution alpha mask that matches the mask detail resolution difference loss function; The preset cutout model is optimized based on the corresponding first error information, second error information, third error information and fourth error information by utilizing the regression loss function, the synthesis loss function, the Laplace loss function and the mask detail resolution difference to obtain a target cutout model.

6. The method according to claim 5, characterized in that The obtaining first error information between the real alpha mask corresponding to the training sample and the target high-resolution alpha mask that matches the regression loss function includes: Acquire first mask information at the to-be-predicted region in the real alpha mask corresponding to the training sample, and second mask information at the to-be-predicted region in the target high-resolution alpha mask; The mean absolute error between the first mask information and the second mask information is obtained as first error information matching the regression loss function.

7. The method according to claim 5, characterized in that The obtaining second error information between the real alpha mask and the target high-resolution alpha mask that matches the synthesis loss function includes: Acquire a first image synthesized by the real alpha mask on the foreground and the background, and a second image synthesized by the target high-resolution alpha mask on the foreground and the background; An absolute difference between the first image and the second image is obtained as second error information matching the synthetic loss function.

8. The method according to claim 5, characterized in that The obtaining third error information between the real alpha mask and the target high-resolution alpha mask that matches the Laplace loss function includes: A Laplace difference between the true alpha mask and the target high-resolution alpha mask is obtained as third error information matched with the Laplace loss function.

9. The method according to claim 5, characterized in that The obtaining fourth error information between the real alpha mask and the target high-resolution alpha mask that matches the mask detail resolution difference loss function includes: Obtaining a mask detail difference between the real alpha mask and the low-resolution alpha mask; Fourth error information matching the mask detail resolution difference loss function is obtained based on the mask detail difference and the detail difference feature output by the detail difference feature extractor.

10. An image cutout device based on alpha mask refinement, characterized in that: include: An acquisition unit, used for acquiring a plurality of training samples, wherein the training samples are high-resolution images; A cutout unit, used for inputting the training sample into a preset cutout model to obtain a low-resolution alpha mask corresponding to the training sample; A feature unit, configured to obtain detail difference features between the training sample and the corresponding low-resolution image using a detail difference feature extractor; A processing unit, configured to fuse the detail difference feature with the low-resolution alpha mask to obtain a target high-resolution alpha mask; The loss unit is used to obtain the error information between the real alpha mask corresponding to the training sample and the target high-resolution alpha mask, and optimize the preset cutout model based on the error information using a preset loss function to obtain a target cutout model, which is applied to the cutout of a large number of high-resolution images.