Image processing model training method and device and image processing method and device
By generating and using image processing models of target, reference and background masks, the problem of being unable to edit each instance in the image separately in the prior art is solved, and fine processing and cutting of complex relationships between instances are realized.
Patent Information
- Application Number
- CN202210550533.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-05-18
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2042-05-18
AI Technical Summary
The existing cutout technology cannot implement separate editing of each instance in the image in many scenarios, especially when there are occlusions or close together between instances, it is impossible to clearly distinguish each instance.
By obtaining the mask for each instance and generating the target mask, reference mask and background mask based on the image processing model, the image processing model is input to predict the transparency of each instance, thereby achieving a combination of instance segmentation and instance extraction.
This method can more accurately predict the transparency of instances in the image, solve the problem of distinction between instances when occluding or close together, and realize fine example cutouts.
Smart Images

Figure CN114863216B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of image processing, and in particular to a method and device for training an image processing model and an image processing method and device. Background Art
[0002] The rapid development of mobile Internet technology has brought more possibilities to image processing. The current Internet platforms provide a variety of image processing methods, such as beautifying and editing uploaded photos, or re-creating images. As a basic image processing and editing technology, cutout technology has achieved great performance improvements under the widespread application of deep neural networks. However, the existing cutout technology still has certain limitations in many scenarios, such as the inability to edit each instance in the image separately. Summary of the invention
[0003] The present disclosure provides a training method and device for an image processing model and an image processing method and device to at least solve the above-mentioned problems in the related art, or may not solve any of the above-mentioned problems. The technical solution of the present disclosure is as follows:
[0004] According to a first aspect of an embodiment of the present disclosure, a training method for an image processing model is provided, comprising: obtaining a training image, wherein the training image contains at least one instance, each instance being marked with a true transparency; obtaining a mask of each instance of the at least one instance; for each instance, based on the mask of the at least one instance, generating a target mask, a reference mask and a background mask of the current instance, and inputting the training image, the target mask, the reference mask and the background mask into the image processing model to obtain the target transparency, the reference transparency and the background transparency of the current instance, wherein the target mask represents the mask of the current instance, the reference mask represents the masks of the remaining instances except the current instance, the background mask represents the mask of the remaining parts of the training image except the at least one instance, the target transparency is the transparency corresponding to the target mask, the reference transparency is the transparency corresponding to the reference mask, and the background transparency is the transparency corresponding to the background mask; calculating a loss based on the target transparency, the reference transparency, the background transparency and the true transparency of each instance; and training the image processing model by adjusting the parameters of the image processing model based on the loss.
[0005] Optionally, the image processing model includes a feature extraction module and a transparency prediction module; the step of inputting the training image, the target mask, the reference mask and the background mask into the image processing model to obtain the target transparency, reference transparency and background transparency of the current instance includes: inputting the training image, the target mask, the reference mask and the background mask into the feature extraction module to obtain image features corresponding to the current instance; and inputting the image features into the transparency prediction module to obtain the target transparency, reference transparency and background transparency of the current instance.
[0006] Optionally, it also includes: for each instance, obtaining the target transparency, reference transparency and background transparency of the remaining instances, and optimizing the target transparency, reference transparency and background transparency of the current instance based on the target transparency, reference transparency and background transparency of the remaining instances to obtain the optimized target transparency, reference transparency and background transparency of the current instance.
[0007] Optionally, the loss is calculated based on the target transparency, the reference transparency, the background transparency and the true transparency of each instance, including: calculating a first loss based on the target transparency, the reference transparency, the background transparency and the true transparency of each instance; calculating a second loss based on the optimized target transparency, the reference transparency, the background transparency and the true transparency of each instance; and determining the loss based on the first loss and the second loss.
[0008] Optionally, the target transparency, reference transparency and background transparency of the current instance are optimized based on the target transparency, reference transparency and background transparency of the remaining instances to obtain the optimized target transparency, reference transparency and background transparency of the current instance, including: for each instance, the following operations are performed: obtaining image features corresponding to at least one instance; optimizing the image features corresponding to the current instance based on the image features corresponding to the remaining instances to obtain the optimized image features corresponding to the current instance; inputting the optimized image features corresponding to the current instance into the transparency prediction module to obtain the optimized target transparency, reference transparency and background transparency corresponding to the current instance.
[0009] Optionally, the optimizing the image features corresponding to the current instance based on the image features corresponding to the remaining instances to obtain the optimized image features corresponding to the current instance includes: decomposing the image features corresponding to each instance of the at least one instance into target sub-features, reference sub-features and background sub-features; performing feature fusion on the target sub-features, reference sub-features and background sub-features corresponding to the current instance based on the target sub-features, reference sub-features and background sub-features corresponding to the remaining instances to obtain the optimized target sub-features, reference sub-features and background sub-features corresponding to the current instance; and aggregating the optimized target sub-features, reference sub-features and background sub-features corresponding to the current instance to obtain the optimized image features corresponding to the current instance.
[0010] Optionally, decomposing the image features corresponding to each instance of the at least one instance into target sub-features, reference sub-features and background sub-features includes: combining the image features corresponding to each instance with the target transparency of each instance to obtain the target sub-features of each instance; combining the image features corresponding to each instance with the reference transparency of each instance to obtain the reference sub-features of each instance; combining the image features corresponding to each instance with the background transparency of each instance to obtain the background sub-features of each instance.
[0011] Optionally, the feature fusion of the target sub-features, reference sub-features and background sub-features corresponding to the current instance based on the target sub-features, reference sub-features and background sub-features corresponding to the remaining instances includes: based on the target sub-features and reference sub-features corresponding to each instance, feature reduction is performed on the target sub-features corresponding to the current instance to obtain an optimized target sub-feature corresponding to the current instance; based on the reference sub-features corresponding to the current instance and the target sub-features corresponding to the remaining instances, feature reduction is performed on the reference sub-features corresponding to the current instance to obtain an optimized reference sub-feature corresponding to the current instance; based on the background sub-features corresponding to each instance, feature reduction is performed on the background sub-features corresponding to the current instance to obtain an optimized background sub-feature corresponding to the current instance.
[0012] Optionally, the loss is calculated based on the target transparency, the reference transparency, the background transparency and the true transparency of each instance, including: determining the first minimum absolute deviation between the target transparency, the reference transparency, the background transparency and the true transparency of each instance; applying smoothness constraints to the parameters of the image processing model based on the target transparency, the reference transparency, the background transparency and the true transparency of each instance to obtain the first Laplace loss of each instance; obtaining the first synthetic constraint loss of each instance based on the target synthesis part, the reference synthesis part, the background synthesis part and the training image of each instance, wherein , the target synthesis part is obtained based on the target transparency and the target instance foreground image, the reference synthesis part is obtained based on the reference transparency and the reference instance foreground image, the background synthesis part is obtained based on the background transparency and the background image, the target instance foreground image, the reference instance foreground image and the background image are all determined based on the real transparency and the training image; based on the target transparency, the reference transparency and the background transparency of each instance, the first transparency constraint loss of each instance is obtained; based on the first minimum absolute deviation, the first Laplace loss, the first synthesis constraint loss and the first transparency constraint loss of each instance, the loss is determined.
[0013] Optionally, the loss is calculated based on the target transparency, the reference transparency, the background transparency and the true transparency of each instance, including: determining a first minimum absolute deviation between the target transparency, the reference transparency, the background transparency and the true transparency of each instance, and determining a second minimum absolute deviation between the optimized target transparency, the reference transparency, the background transparency and the true transparency of each instance; applying smoothness constraints to the parameters of the image processing model based on the target transparency, the reference transparency, the background transparency and the true transparency of each instance to obtain a first Laplace loss for each instance, and applying smoothness constraints to the parameters of the image processing model based on the optimized target transparency, the reference transparency, the background transparency and the true transparency of each instance to obtain a second Laplace loss for each instance; obtaining a first synthesis constraint loss for each instance based on a target synthesis part, a reference synthesis part, a background synthesis part and the training image of each instance, and obtaining a second synthesis constraint loss for each instance based on an optimized target synthesis part, an optimized reference synthesis part, an optimized background synthesis part and the training image of each instance, wherein the target synthesis part is based on the target transparency and the foreground of the target instance. The reference synthesis part is obtained based on the reference transparency and the reference instance foreground image, the background synthesis part is obtained based on the background transparency and the background image, the target instance foreground image, the reference instance foreground image and the background image are all determined based on the real transparency and the training image, the optimized target synthesis part is obtained based on the optimized target transparency and the target instance foreground image, the optimized reference synthesis part is obtained based on the optimized reference transparency and the reference instance foreground image, and the optimized background synthesis part is obtained based on the optimized background transparency and the background image; based on the target transparency, the reference transparency and the background transparency of each instance, a first transparency constraint loss of each instance is obtained, and based on the optimized target transparency, the reference transparency and the background transparency of each instance, a second transparency constraint loss of each instance is obtained; based on the first minimum absolute deviation, the first Laplace loss, the first synthesis constraint loss and the first transparency constraint loss of each instance, the first loss is determined, and based on the second minimum absolute deviation, the second Laplace loss, the second synthesis constraint loss and the second transparency constraint loss of each instance, the second loss is determined; based on the first loss and the second loss, the loss is determined.
[0014] According to a second aspect of an embodiment of the present disclosure, there is provided an image processing method, comprising: obtaining an image to be processed, wherein the image to be processed contains at least one instance; obtaining a mask of each instance of the at least one instance; for each instance, based on the mask of the at least one instance, generating a target mask, a reference mask and a background mask of the current instance, and inputting the image to be processed, the target mask, the reference mask and the background mask into an image processing model to obtain a target transparency, a reference transparency and a background transparency of the current instance, wherein the target mask represents the mask of the current instance, the reference mask represents the masks of the remaining instances except the current instance, and the background mask represents the mask of the remaining parts of the image to be processed except the at least one instance, the target transparency is the transparency corresponding to the target mask, the reference transparency is the transparency corresponding to the reference mask, and the background transparency is the transparency corresponding to the background mask; based on the target transparency, the reference transparency and the background transparency of each instance, extracting each instance from the image to be processed.
[0015] Optionally, the image processing model includes a feature extraction module and a transparency prediction module; the step of inputting the image to be processed, the target mask, the reference mask and the background mask into the image processing model to obtain the target transparency, reference transparency and background transparency of the current instance includes: inputting the image to be processed, the target mask, the reference mask and the background mask into the feature extraction module to obtain image features corresponding to the current instance; and inputting the image features into the transparency prediction module to obtain the target transparency, reference transparency and background transparency of the current instance.
[0016] Optionally, it also includes: for each instance, obtaining the target transparency, reference transparency and background transparency of the remaining instances, and optimizing the target transparency, reference transparency and background transparency of the current instance based on the target transparency, reference transparency and background transparency of the remaining instances to obtain the optimized target transparency, reference transparency and background transparency of the current instance.
[0017] Optionally, extracting each instance from the image to be processed based on the target transparency, the reference transparency and the background transparency of each instance includes: extracting each instance from the image to be processed based on the optimized target transparency, reference transparency and background transparency of each instance.
[0018] Optionally, the target transparency, reference transparency and background transparency of the current instance are optimized based on the target transparency, reference transparency and background transparency of the remaining instances to obtain the optimized target transparency, reference transparency and background transparency of the current instance, including: for each instance, the following operations are performed: obtaining image features corresponding to at least one instance; optimizing the image features corresponding to the current instance based on the image features corresponding to the remaining instances to obtain the optimized image features corresponding to the current instance; inputting the optimized image features corresponding to the current instance into the transparency prediction module to obtain the optimized target transparency, reference transparency and background transparency corresponding to the current instance.
[0019] Optionally, the optimizing the image features corresponding to the current instance based on the image features corresponding to the remaining instances to obtain the optimized image features corresponding to the current instance includes: decomposing the image features corresponding to each instance of the at least one instance into target sub-features, reference sub-features and background sub-features; performing feature fusion on the target sub-features, reference sub-features and background sub-features corresponding to the current instance based on the target sub-features, reference sub-features and background sub-features corresponding to the remaining instances to obtain the optimized target sub-features, reference sub-features and background sub-features corresponding to the current instance; and aggregating the optimized target sub-features, reference sub-features and background sub-features corresponding to the current instance to obtain the optimized image features corresponding to the current instance.
[0020] Optionally, decomposing the image features corresponding to each instance of the at least one instance into target sub-features, reference sub-features and background sub-features includes: combining the image features corresponding to each instance with the target transparency of each instance to obtain the target sub-features of each instance; combining the image features corresponding to each instance with the reference transparency of each instance to obtain the reference sub-features of each instance; combining the image features corresponding to each instance with the background transparency of each instance to obtain the background sub-features of each instance.
[0021] Optionally, the feature fusion of the target sub-features, reference sub-features and background sub-features corresponding to the current instance based on the target sub-features, reference sub-features and background sub-features corresponding to the remaining instances includes: based on the target sub-features and reference sub-features corresponding to each instance, feature reduction is performed on the target sub-features corresponding to the current instance to obtain an optimized target sub-feature corresponding to the current instance; based on the reference sub-features corresponding to the current instance and the target sub-features corresponding to the remaining instances, feature reduction is performed on the reference sub-features corresponding to the current instance to obtain an optimized reference sub-feature corresponding to the current instance; based on the background sub-features corresponding to each instance, feature reduction is performed on the background sub-features corresponding to the current instance to obtain an optimized background sub-feature corresponding to the current instance.
[0022] According to a third aspect of an embodiment of the present disclosure, a training device for an image processing model is provided, comprising: a first image acquisition unit, configured to: acquire a training image, wherein the training image contains at least one instance, and each instance is marked with a true transparency; a first mask acquisition unit, configured to: acquire a mask of each instance of the at least one instance; a first model prediction unit, configured to: for each instance, based on the mask of the at least one instance, generate a target mask, a reference mask, and a background mask of the current instance, and input the training image, the target mask, the reference mask, and the background mask into the image processing model to obtain the target transparency, the reference transparency, and the background transparency of the current instance, wherein the target The mask represents the mask of the current instance, the reference mask represents the masks of the remaining instances except the current instance, the background mask represents the mask of the remaining parts except the at least one instance in the training image, the target transparency is the transparency corresponding to the target mask, the reference transparency is the transparency corresponding to the reference mask, and the background transparency is the transparency corresponding to the background mask; the loss calculation unit is configured to: calculate the loss based on the target transparency, the reference transparency, the background transparency and the true transparency of each instance; the parameter adjustment unit is configured to: train the image processing model by adjusting the parameters of the image processing model based on the loss.
[0023] Optionally, the image processing model includes a feature extraction module and a transparency prediction module; the first model prediction unit is configured to: input the training image, the target mask, the reference mask and the background mask into the feature extraction module to obtain the image features corresponding to the current instance; input the image features into the transparency prediction module to obtain the target transparency, reference transparency and background transparency of the current instance.
[0024] Optionally, it also includes a first optimization unit, which is configured to: for each instance, obtain the target transparency, reference transparency and background transparency of the remaining instances, and optimize the target transparency, reference transparency and background transparency of the current instance based on the target transparency, reference transparency and background transparency of the remaining instances to obtain the optimized target transparency, reference transparency and background transparency of the current instance.
[0025] Optionally, the loss calculation unit is configured to: calculate a first loss based on the target transparency, the reference transparency, the background transparency and the true transparency of each instance; calculate a second loss based on the optimized target transparency, the reference transparency, the background transparency and the true transparency of each instance; and determine the loss based on the first loss and the second loss.
[0026] Optionally, the first optimization unit is configured to: for each instance, perform the following operations: obtain image features corresponding to at least one instance; optimize the image features corresponding to the current instance based on the image features corresponding to the remaining instances to obtain optimized image features corresponding to the current instance; input the optimized image features corresponding to the current instance into the transparency prediction module to obtain optimized target transparency, reference transparency and background transparency corresponding to the current instance.
[0027] Optionally, the first optimization unit is configured to: decompose the image features corresponding to each instance of the at least one instance into target sub-features, reference sub-features and background sub-features; perform feature fusion on the target sub-features, reference sub-features and background sub-features corresponding to the current instance based on the target sub-features, reference sub-features and background sub-features corresponding to the remaining instances to obtain optimized target sub-features, reference sub-features and background sub-features corresponding to the current instance; and aggregate the optimized target sub-features, reference sub-features and background sub-features corresponding to the current instance to obtain optimized image features corresponding to the current instance.
[0028] Optionally, the first optimization unit is configured to: combine the image features corresponding to each instance with the target transparency of each instance to obtain the target sub-features of each instance; combine the image features corresponding to each instance with the reference transparency of each instance to obtain the reference sub-features of each instance; combine the image features corresponding to each instance with the background transparency of each instance to obtain the background sub-features of each instance.
[0029] Optionally, the first optimization unit is configured to: perform feature reduction on the target sub-feature corresponding to the current instance based on the target sub-feature and the reference sub-feature corresponding to each instance to obtain the optimized target sub-feature corresponding to the current instance; perform feature reduction on the reference sub-feature corresponding to the current instance based on the reference sub-feature corresponding to the current instance and the target sub-features corresponding to the remaining instances to obtain the optimized reference sub-feature corresponding to the current instance; perform feature reduction on the background sub-feature corresponding to the current instance based on the background sub-feature corresponding to each instance to obtain the optimized background sub-feature corresponding to the current instance.
[0030] Optionally, the loss calculation unit is configured to: determine the first minimum absolute deviation between the target transparency, the reference transparency, the background transparency, and the true transparency of each instance; based on the target transparency, the reference transparency, the background transparency and the true transparency of each instance, apply a smoothness constraint to the parameters of the image processing model to obtain the first Laplace loss of each instance; based on the target synthesis part, the reference synthesis part, the background synthesis part and the training image of each instance, obtain the first synthesis constraint loss of each instance, wherein the target synthesis part is obtained based on the target transparency and the target instance foreground image, the reference synthesis part is obtained based on the reference transparency and the reference instance foreground image, the background synthesis part is obtained based on the background transparency and the background image, and the target instance foreground image, the reference instance foreground image and the background image are all determined based on the true transparency and the training image; based on the target transparency, the reference transparency and the background transparency of each instance, obtain the first transparency constraint loss of each instance; based on the first minimum absolute deviation, the first Laplace loss, the first synthesis constraint loss and the first transparency constraint loss of each instance, determine the loss.
[0031] Optionally, the loss calculation unit is configured to: determine a first minimum absolute deviation between the target transparency, the reference transparency, the background transparency, and the true transparency of each instance, and determine a second minimum absolute deviation between the optimized target transparency, the reference transparency, the background transparency, and the true transparency of each instance; based on the target transparency, the reference transparency, the background transparency, and the true transparency of each instance, apply a smoothness constraint to the parameters of the image processing model to obtain a first Laplace loss for each instance, and based on the optimized target transparency, the reference transparency, the background transparency, and the true transparency of each instance, apply a smoothness constraint to the parameters of the image processing model to obtain a second Laplace loss for each instance; based on the target synthesis part, the reference synthesis part, the background synthesis part, and the training image of each instance, obtain a first synthesis constraint loss for each instance, and based on the optimized target synthesis part, the optimized reference synthesis part, the optimized background synthesis part, and the training image of each instance, obtain a second synthesis constraint loss for each instance, wherein the target synthesis part is obtained based on the target transparency and the target instance foreground image, and the reference synthesis part is obtained based on the reference transparency. The target instance foreground image, the reference instance foreground image and the background image are all determined based on the real transparency and the training image, the optimized target synthesis part is obtained based on the optimized target transparency and the target instance foreground image, the optimized reference synthesis part is obtained based on the optimized reference transparency and the reference instance foreground image, and the optimized background synthesis part is obtained based on the optimized background transparency and the background image; based on the target transparency, the reference transparency and the background transparency of each instance, a first transparency constraint loss of each instance is obtained, and based on the optimized target transparency, the reference transparency and the background transparency of each instance, a second transparency constraint loss of each instance is obtained; based on the first minimum absolute deviation, the first Laplace loss, the first synthesis constraint loss and the first transparency constraint loss of each instance, the first loss is determined, and based on the second minimum absolute deviation, the second Laplace loss, the second synthesis constraint loss and the second transparency constraint loss of each instance, the second loss is determined; based on the first loss and the second loss, the loss is determined.
[0032] According to a fourth aspect of an embodiment of the present disclosure, an image processing device is provided, comprising: a second image acquisition unit, configured to: acquire an image to be processed, wherein the image to be processed includes at least one instance; a second mask acquisition unit, configured to: acquire a mask of each instance of the at least one instance; a second model prediction unit, configured to: for each instance, based on the mask of the at least one instance, generate a target mask, a reference mask and a background mask of the current instance, and input the image to be processed, the target mask, the reference mask and the background mask into an image processing model to obtain a target transparency, a reference transparency and a background transparency of the current instance. and background transparency, wherein the target mask represents the mask of the current instance, the reference mask represents the masks of the remaining instances except the current instance, the background mask represents the masks of the remaining parts of the image to be processed except the at least one instance, the target transparency is the transparency corresponding to the target mask, the reference transparency is the transparency corresponding to the reference mask, and the background transparency is the transparency corresponding to the background mask; an instance extraction unit is configured to extract each instance from the image to be processed based on the target transparency, the reference transparency and the background transparency of each instance.
[0033] Optionally, the image processing model includes a feature extraction module and a transparency prediction module; the second model prediction unit is configured to: input the image to be processed, the target mask, the reference mask and the background mask into the feature extraction module to obtain the image features corresponding to the current instance; input the image features into the transparency prediction module to obtain the target transparency, reference transparency and background transparency of the current instance.
[0034] Optionally, a second optimization unit is also included, which is configured to: for each instance, obtain the target transparency, reference transparency and background transparency of the remaining instances, and optimize the target transparency, reference transparency and background transparency of the current instance based on the target transparency, reference transparency and background transparency of the remaining instances to obtain the optimized target transparency, reference transparency and background transparency of the current instance.
[0035] Optionally, the instance extraction unit is configured to extract each instance from the image to be processed based on the optimized target transparency, reference transparency and background transparency of each instance.
[0036] Optionally, the second optimization unit is configured to: for each instance, perform the following operations: obtain image features corresponding to at least one instance; optimize the image features corresponding to the current instance based on the image features corresponding to the remaining instances to obtain optimized image features corresponding to the current instance; input the optimized image features corresponding to the current instance into the transparency prediction module to obtain optimized target transparency, reference transparency and background transparency corresponding to the current instance.
[0037] Optionally, the second optimization unit is configured to: decompose the image features corresponding to each instance of the at least one instance into target sub-features, reference sub-features and background sub-features; perform feature fusion on the target sub-features, reference sub-features and background sub-features corresponding to the current instance based on the target sub-features, reference sub-features and background sub-features corresponding to the remaining instances to obtain optimized target sub-features, reference sub-features and background sub-features corresponding to the current instance; and aggregate the optimized target sub-features, reference sub-features and background sub-features corresponding to the current instance to obtain optimized image features corresponding to the current instance.
[0038] Optionally, the second optimization unit is configured to: combine the image features corresponding to each instance with the target transparency of each instance to obtain the target sub-features of each instance; combine the image features corresponding to each instance with the reference transparency of each instance to obtain the reference sub-features of each instance; combine the image features corresponding to each instance with the background transparency of each instance to obtain the background sub-features of each instance.
[0039] Optionally, the second optimization unit is configured to: perform feature reduction on the target sub-feature corresponding to the current instance based on the target sub-feature and the reference sub-feature corresponding to each instance to obtain the optimized target sub-feature corresponding to the current instance; perform feature reduction on the reference sub-feature corresponding to the current instance based on the reference sub-feature corresponding to the current instance and the target sub-features corresponding to the remaining instances to obtain the optimized reference sub-feature corresponding to the current instance; perform feature reduction on the background sub-feature corresponding to the current instance based on the background sub-feature corresponding to each instance to obtain the optimized background sub-feature corresponding to the current instance.
[0040] According to a fifth aspect of an embodiment of the present disclosure, there is provided an electronic device, comprising: at least one processor; and at least one memory storing computer executable instructions, wherein when the computer executable instructions are executed by the at least one processor, the at least one processor is prompted to execute a training method for an image processing model or an image processing method according to the present disclosure.
[0041] According to a sixth aspect of an embodiment of the present disclosure, a computer-readable storage medium storing instructions is provided. When the instructions are executed by at least one processor, the at least one processor is prompted to execute a training method for an image processing model or an image processing method according to the present disclosure.
[0042] According to a seventh aspect of an embodiment of the present disclosure, a computer program product is provided, comprising computer instructions, which, when executed by at least one processor, implement the training method of the image processing model or the image processing method according to the present disclosure.
[0043] The technical solution provided by the embodiments of the present disclosure brings at least the following beneficial effects:
[0044] According to the training method and device of the image processing model and the image processing method and device disclosed in the present invention, by obtaining the mask of each instance and obtaining the transparency of each instance based on the image processing model, instance segmentation and instance extraction can be combined. By generating the target mask, reference mask and background mask of the instance, the image processing model can predict the target transparency, reference transparency and background transparency, so that the image processing model can not only learn the difference between the instance and the background, but also learn the difference between the instances, so as to obtain more accurate prediction results, and can solve the problem in the prior art that when the instances are blocked or close together, the instances cannot be clearly distinguished, and complete the fine instance cutout.
[0045] In addition, according to the training method and device of the image processing model and the image processing method and device disclosed in the present invention, instances can be optimized, and information interaction can be performed between instances to obtain more accurate and precise prediction results.
[0046] In addition, the training method and device of the image processing model and the image processing method and device according to the present invention can be easily applied to many practical scenarios, such as extracting only the image of the speaker in a meeting scene and ignoring other people in the background, or realizing personalized editing for individuals in picture editing applications.
[0047] It is to be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS
[0048] The drawings herein are incorporated into and constitute a part of the specification, illustrate embodiments consistent with the present disclosure, and together with the description are used to explain the principles of the present disclosure, and do not constitute improper limitations on the present disclosure.
[0049] Figure 1 The figure is a flowchart of a method for training an image processing model according to an exemplary embodiment.
[0050] Figure 2 It is a framework diagram of an image processing model according to an exemplary embodiment.
[0051] Figure 3 The diagram is a qualitative test comparison diagram according to an exemplary embodiment.
[0052] Figure 4 The figure is a flowchart of an image processing method according to an exemplary embodiment.
[0053] Figure 5 It is a block diagram of a training device for an image processing model according to an exemplary embodiment.
[0054] Figure 6 is a block diagram of an image processing apparatus according to an exemplary embodiment.
[0055] Figure 7 is a block diagram of an electronic device 700 according to an exemplary embodiment. DETAILED DESCRIPTION
[0056] In order to enable ordinary persons in the art to better understand the technical solutions of the present disclosure, the technical solutions in the embodiments of the present disclosure will be clearly and completely described below in conjunction with the accompanying drawings.
[0057] It should be noted that the terms "first", "second", etc. in the specification and claims of the present disclosure and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged where appropriate, so that the embodiments of the present disclosure described herein can be implemented in an order other than those illustrated or described herein. The implementation methods described in the following examples do not represent all implementation methods consistent with the present disclosure. Instead, they are merely examples of devices and methods consistent with some aspects of the present disclosure as detailed in the attached claims.
[0058] It should be noted that the phrase "at least one of the items" in the present disclosure includes three types of parallel situations: "any one of the items", "a combination of any number of the items", and "all of the items". For example, "including at least one of A and B" includes the following three parallel situations: (1) including A; (2) including B; (3) including A and B. Another example is "executing at least one of step 1 and step 2" which means the following three parallel situations: (1) executing step 1; (2) executing step 2; (3) executing step 1 and step 2.
[0059] The rapid development of mobile Internet technology has brought more possibilities to image processing. The current Internet platforms provide a variety of image processing methods. For example, on self-media platforms, uploaded photos can be beautified and edited, or images can be re-created. As a basic image processing and editing technology, cutout technology has achieved great performance improvements with the widespread application of deep neural networks. However, the existing cutout technology still has certain limitations in many scenarios, such as the inability to edit each instance in the image separately.
[0060] Although existing technologies such as instance segmentation algorithms can distinguish individual instances, the segmentation results in a relatively coarse binary mask and cannot handle fine structures such as hair. Most related technologies use a cutout algorithm to cut out a single instance scene. However, although the existing cutout algorithms can extract fine instance foregrounds, they cannot distinguish individual instances.
[0061] For the existing instance segmentation algorithm and cutout algorithm, a feasible solution is to make a simple combination of the instance segmentation algorithm and the cutout algorithm, that is, given an image, first use the existing instance segmentation algorithm to detect the instance individuals in the image and obtain the mask; then use the mask to generate the trimap, combine it with the existing trimap-based cutout algorithm, extract the transparency alpha of each instance individual in the image, and then obtain the instance individual to be cut out. However, this method has two shortcomings. One is that the accuracy of the cutout will be affected by the trimap generated by the mask. When the trimap cannot reasonably cover the foreground, background and unknown area, the cutout instance individual will have a large deviation. The second is that this method does not take into account the information interaction between instance individuals. When there is occlusion between instances or they are close together, it is not possible to clearly distinguish each instance.
[0062] In order to solve the problems existing in the above-mentioned related technologies, the present disclosure proposes a training method and device for an image processing model and an image processing method and device, which obtains the mask of each instance and obtains the transparency of each instance based on the image processing model, so that instance segmentation and instance extraction can be combined. By generating the target mask, reference mask and background mask of the instance, the image processing model can predict the target transparency, reference transparency and background transparency, so that the image processing model can not only learn the difference between the instance and the background, but also learn the difference between the instances, so as to obtain more accurate prediction results, which can solve the problem in the prior art that when the instances are blocked or close together, the instances cannot be clearly distinguished, and complete the fine instance cutout.
[0063] Next, we will refer to Figures 1 to 7The training method and device of the image processing model and the image processing method and device according to the present disclosure are described in detail.
[0064] Figure 1 FIG. 1 is a flowchart of a method for training an image processing model according to an exemplary embodiment. Figure 1 In step 101, a training image may be obtained, wherein the training image includes at least one instance, and each instance is marked with true transparency.
[0065] According to an exemplary embodiment of the present disclosure, in the process of training the image processing model, the training images used may include many, Figure 1 The process shown is described with only one training image. For all images used to train the image processing model, the same process can be used. Figure 1 The method flow shown is not repeated here. The training image can be, but is not limited to, an image containing at least one instance, and the instance can be an individual person or an individual object for extraction and subsequent editing processing (i.e., cutout). For the exemplary embodiments of the present disclosure, for each instance of each training image, a real transparency is marked, and the real transparency can be used as an accurate real value for training to adjust the parameters of the model in combination with the predicted value of the model to complete the training of the model.
[0066] It should be noted here that the transparency in the present disclosure may be, but is not limited to, an alpha channel value (alpha).
[0067] An exemplary embodiment of the present disclosure may first perform instance segmentation on at least one instance in a training image, and then perform subsequent transparency prediction on each instance. The following describes the instance segmentation and transparency prediction in steps 102 and 103, respectively.
[0068] At step 102 , a mask for each instance of at least one instance may be obtained.
[0069] According to an exemplary embodiment of the present disclosure, the mask of each instance can be obtained through instance segmentation to perform instance segmentation, which can be implemented through existing deep learning methods. For example, the training image can be input into a trained instance segmentation model to obtain the mask of each instance. Here, the trained instance segmentation model can be an existing structure such as MaskRCNN, or other models that can perform instance segmentation, and the present disclosure does not limit this.
[0070] In step 103, for each instance, based on the mask of at least one instance, a target mask, a reference mask and a background mask of the current instance may be generated, and the training image, the target mask, the reference mask and the background mask may be input into the image processing model to obtain the target transparency, the reference transparency and the background transparency of the current instance, wherein the target mask represents the mask of the current instance, the reference mask represents the masks of the remaining instances except the current instance, the background mask represents the mask of the remaining parts except at least one instance in the training image, the target transparency is the transparency corresponding to the target mask, the reference transparency is the transparency corresponding to the reference mask, and the background transparency is the transparency corresponding to the background mask.
[0071] According to an exemplary embodiment of the present disclosure, a correspondence between each instance and a training image can be established. Then, from the perspective of the instance, the training image can be distinguished into the cost instance, the remaining instances other than the instance, and the image portion other than the instance. That is, the entire training image can be distinguished into the target instance (current instance), the reference instance (the remaining instances other than the current instance), and the background (the remaining portion of the training image other than the instance). Based on the above analysis, an exemplary embodiment of the present disclosure can generate a tri-mask of the current instance based on a mask of at least one instance, and the tri-mask includes a target mask, a reference mask, and a background mask.
[0072] For example, for the i-th instance of the training image, three masks can be generated by the following equations (1) to (3):
[0073] M i,t =M i (1)
[0074]
[0075] M i,b =1-M i,t ∪M i,r (3)
[0076] Among them, M i,t is the target mask of the i-th instance, M i is the mask of the i-th instance, M i,r is the reference mask of the i-th instance, n is the number of instances in the training image, i∈{1,2,…,n}, M j is the mask of the jth instance, M i,b is the background mask of the i-th instance.
[0077] Based on the above description, the above three masks and training images can be input into the image processing model to predict transparency and obtain tri-matte, which includes target transparency, reference transparency and background transparency, corresponding to the target mask, reference mask and background mask respectively. The image processing model in the exemplary embodiment of the present disclosure can be, but not limited to, U-Net, ResNet34, ResNet50, ResNet101, MobileNet series. For example, when higher image processing accuracy is required, ResNet50 or ResNet101 can be used, and when faster processing speed is required, MobileNet series can be used. In the case where the image processing model is U-Net, it includes an encoder for encoding input items into features of different scales, and then based on a decoder, it reconstructs the original scale layer by layer from small to large according to these features, and finally sends the reconstructed features to a transparency prediction head composed of convolutional layers for prediction.
[0078] The image processing model in the exemplary embodiment of the present disclosure may include a feature extraction module and a transparency prediction module. Then, with respect to the structure of the image processing model, when predicting each transparency, the prediction may be carried out in a division of labor and cooperation among the modules. Specifically, the training image, target mask, reference mask and background mask are first input into the feature extraction module to obtain the image features corresponding to the current instance. Then, the image features are input into the transparency prediction module to obtain the target transparency, reference transparency and background transparency of the current instance. It should be noted here that the image features may be the features generated by the last layer of the image processing model.
[0079] Back to Figure 1 In step 104, the loss may be calculated based on the target transparency, the reference transparency, the background transparency, and the true transparency of each instance.
[0080] According to an exemplary embodiment of the present disclosure, for each instance of a training image, the loss may include four parts, namely, the first minimum absolute deviation, the first Laplace loss, the first synthesis constraint loss and the first transparency constraint loss. Then the loss may be the sum of the above four losses of all instances. The specific process of calculating the loss is specifically explained below.
[0081] First, the first minimum absolute deviation between the target transparency, reference transparency, background transparency and true transparency of each instance can be determined. It should be noted that the true transparency in the exemplary embodiments of the present disclosure may include true target transparency, true reference transparency and true background transparency. The minimum absolute deviation of the present disclosure may be L1 loss.
[0082] Then, based on the target transparency, reference transparency, background transparency and true transparency of each instance, smoothness constraints may be imposed on the parameters of the image processing model to obtain the first Laplacian loss of each instance.
[0083] Next, the first synthesis constraint loss of each instance can be obtained based on the target synthesis part, reference synthesis part, background synthesis part and training image of each instance, wherein the target synthesis part is obtained based on the target transparency and the target instance foreground image, the reference synthesis part is obtained based on the reference transparency and the reference instance foreground image, the background synthesis part is obtained based on the background transparency and the background image, and the target instance foreground image, the reference instance foreground image and the background image are all determined based on the true transparency and the training image.
[0084] For example, the first synthetic constraint loss can be obtained by the following formula (4):
[0085] L mc =||α t F t +α r F r +α b F b -I||1 (4)
[0086] Among them, L mc is the first synthetic constraint loss, α t F t is the target synthesis part, α t is the target transparency, F t is the target instance foreground image, α r F r For reference synthesis, α r is the reference transparency, F r is the reference instance foreground image, α b F b is the background synthesis part, α b is the background transparency, F b is the background image, I is the training image, ‖α t F t +α r F r +α b F b -I‖1 indicates the t F t +α r F r +α b F b -I finds the L1 norm, F t 、F r and F bIt can be calculated from the real transparency and the training image.
[0087] Furthermore, the first transparency constraint loss of each instance may be obtained based on the target transparency, the reference transparency, and the background transparency of each instance.
[0088] For example, the first transparency constraint loss can be obtained by the following formula (5):
[0089] L mα =||α t +α r +α b -1||1 (5)
[0090] Among them, L mα is the first transparency constraint loss, α t is the target transparency, α r is the reference transparency, α b is the background transparency, ‖α t +α r +α b -1‖1 means that α t +α r +α b -1 finds the L1 norm.
[0091] Finally, the loss may be determined based on the first minimum absolute deviation, the first Laplace loss, the first synthesis constraint loss, and the first transparency constraint loss of each instance. For example, but not limited to, the first minimum absolute deviation, the first Laplace loss, the first synthesis constraint loss, and the first transparency constraint loss of each instance may be added together to obtain the loss.
[0092] It should be noted that the above order of obtaining the four losses of each instance is only illustrative, and other orders may be used for calculation, and the present disclosure does not impose any limitation on this.
[0093] According to the embodiment, through the above-mentioned first synthesis constraint loss and first transparency constraint loss, the loss can be determined for the predicted target transparency, reference transparency and background transparency in a targeted manner.
[0094] For the above prediction process, although the guidance information of the remaining instances and the background has been introduced by inputting the reference mask and the background mask, there is no information interaction between the instances during the prediction process, which may cause the tri-matte predicted by each instance of the same training image to be inconsistent. For example, each instance of the same training image will predict the background transparency. In theory, the background transparency predicted by each instance should be consistent, but due to the difference in input information, the obtained background transparency may be inconsistent. In order to solve the problem of inconsistent predictions caused by information differences, after the transparency prediction module predicts each transparency, instance optimization (multi-instance refinement) can be performed to allow the maximum degree of freedom of information interaction between instances, and the transparency is predicted again based on the result after the interaction to obtain the optimized transparency. This process is described in detail below.
[0095] Specifically, for each instance, the target transparency, reference transparency, and background transparency of the remaining instances can be obtained, and the target transparency, reference transparency, and background transparency of the current instance can be optimized based on the target transparency, reference transparency, and background transparency of the remaining instances to obtain the optimized target transparency, reference transparency, and background transparency of the current instance. According to an embodiment, by optimizing the target transparency, reference transparency, and background transparency of the current instance, and optimizing based on the target transparency, reference transparency, and background transparency of the remaining instances, information interaction can be performed between instances.
[0096] Furthermore, the image features obtained in the feature extraction module can be used to implement the optimization. Then, for each instance, the following operations can be performed: first, the image features corresponding to at least one instance are obtained. Then, the image features corresponding to the current instance are optimized based on the image features corresponding to the remaining instances to obtain the optimized image features corresponding to the current instance. Finally, the optimized image features corresponding to the current instance are input into the transparency prediction module to obtain the optimized target transparency, reference transparency and background transparency corresponding to the current instance. According to the embodiment, the target transparency, reference transparency and background transparency of the current instance can be optimized based on the image features obtained by the feature extraction module.
[0097] Here, the image features corresponding to the current instance can be optimized in the order of feature decomposition, feature propagation, and feature aggregation. For feature propagation, the decomposed features of each instance can be propagated to the remaining instances, so that each instance receives the decomposed features from all the remaining instances. For each training image, the number of instances corresponding to it may be different, so the embodiment of the present disclosure adopts a fusion method to perform feature propagation.
[0098] First, for feature decomposition, image features corresponding to each instance of at least one instance may be decomposed into target sub-features, reference sub-features, and background sub-features.
[0099] Here, the image feature corresponding to each instance can be combined with the target transparency of each instance to obtain the target sub-feature of each instance. The image feature corresponding to each instance and the reference transparency of each instance can be combined to obtain the reference sub-feature of each instance. The image feature corresponding to each instance and the background transparency of each instance can be combined to obtain the background sub-feature of each instance.
[0100] For example, the image feature corresponding to each instance may be multiplied by the target transparency, reference transparency and background transparency corresponding to the instance respectively to obtain the target sub-feature, reference sub-feature and background sub-feature corresponding to the corresponding instance.
[0101] Next, for feature propagation, the target sub-features, reference sub-features, and background sub-features corresponding to the current instance may be feature fused based on the target sub-features, reference sub-features, and background sub-features corresponding to the remaining instances to obtain optimized target sub-features, reference sub-features, and background sub-features corresponding to the current instance. For example, an exemplary embodiment of the present disclosure may perform feature fusion in a feature reduction manner.
[0102] Here, based on the target sub-features and reference sub-features corresponding to each instance, feature reduction may be performed on the target sub-features corresponding to the current instance to obtain an optimized target sub-feature corresponding to the current instance.
[0103] For example, the optimized target sub-feature corresponding to the i-th instance can be obtained by the following formula (6):
[0104]
[0105] in, is the optimized target sub-feature corresponding to the i-th instance, is the target sub-feature of the i-th instance, n is the number of instances, i∈{1,2,…,n}, is the reference sub-feature of the jth instance, is the target sub-feature of the j-th instance.
[0106] Based on the reference sub-feature corresponding to the current instance and the target sub-features corresponding to the remaining instances, feature reduction may be performed on the reference sub-feature corresponding to the current instance to obtain an optimized reference sub-feature corresponding to the current instance.
[0107] For example, the optimized reference sub-feature corresponding to the i-th instance can be obtained by the following formula (7):
[0108]
[0109] in, is the optimized reference sub-feature corresponding to the i-th instance, is the reference sub-feature of the i-th instance, n is the number of instances, i∈{1,2,…,n}, is the target sub-feature of the j-th instance.
[0110] Based on the background sub-features corresponding to each instance, feature reduction may be performed on the background sub-features corresponding to the current instance to obtain optimized background sub-features corresponding to the current instance.
[0111] For example, the optimized background sub-feature corresponding to the i-th instance can be obtained by the following formula (8):
[0112]
[0113] in, is the optimized background sub-feature corresponding to the i-th instance, is the background sub-feature of the jth instance, n is the number of instances, i∈{1,2,…,n}.
[0114] According to the embodiment, the optimized target sub-feature, optimized reference sub-feature and optimized background sub-feature corresponding to the current instance can be obtained by feature reduction, and feature fusion can be achieved when the number of instances of the image input to the image processing model is different.
[0115] Finally, for feature aggregation, the optimized target sub-features, reference sub-features, and background sub-features corresponding to the current instance may be aggregated to obtain the optimized image features corresponding to the current instance. It should be noted that in the embodiments of the present disclosure, the aggregation operation may be implemented through a convolutional layer.
[0116] According to the embodiment, through the process of feature decomposition, feature propagation and feature aggregation, it is possible to optimize the target transparency, reference transparency and background transparency of the current instance based on the image features.
[0117] Based on the above optimization of transparency, the loss can be achieved through the following steps:
[0118] First, a first loss may be calculated based on the target transparency, the reference transparency, the background transparency, and the true transparency of each instance.
[0119] Then, the second loss may be calculated based on the optimized target transparency, reference transparency, background transparency, and true transparency of each instance.
[0120] Finally, the loss may be determined based on the first loss and the second loss.
[0121] According to the embodiment, the target transparency, reference transparency, background transparency of each instance, as well as the optimized target transparency, reference transparency, background transparency of each instance are considered in the process of determining the loss, which can enable the trained image processing model to have more accurate prediction results.
[0122] Based on the above-mentioned loss determination steps, specifically, first, the first minimum absolute deviation between the target transparency, reference transparency, background transparency, and true transparency of each instance can be determined, and the second minimum absolute deviation between the optimized target transparency, reference transparency, background transparency, and true transparency of each instance can be determined.
[0123] Then, based on the target transparency, reference transparency, background transparency and true transparency of each instance, smoothness constraints can be applied to the parameters of the image processing model to obtain the first Laplace loss of each instance, and based on the optimized target transparency, reference transparency, background transparency and true transparency of each instance, smoothness constraints can be applied to the parameters of the image processing model to obtain the second Laplace loss of each instance.
[0124] Next, a first synthesis constraint loss for each instance can be obtained based on the target synthesis part, reference synthesis part, background synthesis part and training image of each instance, and a second synthesis constraint loss for each instance can be obtained based on the optimized target synthesis part, optimized reference synthesis part, optimized background synthesis part and training image of each instance, wherein the target synthesis part is obtained based on the target transparency and the target instance foreground image, the reference synthesis part is obtained based on the reference transparency and the reference instance foreground image, the background synthesis part is obtained based on the background transparency and the background image, the target instance foreground image, the reference instance foreground image and the background image are all determined based on the real transparency and the training image, the optimized target synthesis part is obtained based on the optimized target transparency and the target instance foreground image, the optimized reference synthesis part is obtained based on the optimized reference transparency and the reference instance foreground image, and the optimized background synthesis part is obtained based on the optimized background transparency and the background image.
[0125] Then, the first transparency constraint loss of each instance can be obtained based on the target transparency, reference transparency and background transparency of each instance, and the second transparency constraint loss of each instance can be obtained based on the optimized target transparency, reference transparency and background transparency of each instance.
[0126] Next, the first loss may be determined based on the first minimum absolute deviation, the first Laplace loss, the first synthetic constraint loss, and the first transparency constraint loss of each instance, and the second loss may be determined based on the second minimum absolute deviation, the second Laplace loss, the second synthetic constraint loss, and the second transparency constraint loss of each instance. For example, the first minimum absolute deviation, the first Laplace loss, the first synthetic constraint loss, and the first transparency constraint loss of each instance may be added together to obtain the first loss; the second minimum absolute deviation, the second Laplace loss, the second synthetic constraint loss, and the second transparency constraint loss of each instance may be added together to obtain the second loss.
[0127] Finally, the loss may be determined based on the first loss and the second loss. For example, the first loss and the second loss may be added together to obtain the loss.
[0128] According to the embodiment, through the above-mentioned first synthetic constraint loss and first transparency constraint loss, second synthetic constraint loss and second transparency constraint loss, the loss of the predicted target transparency, reference transparency and background transparency, as well as the optimized target transparency, reference transparency and background transparency can be determined in a targeted manner.
[0129] It should be noted that the order of obtaining the above losses is only illustrative, and other orders may be used for calculation, and the present disclosure does not limit this. The second synthetic constraint loss and the second transparency constraint loss may respectively adopt the above formulas (4) and (5), and the corresponding transparency in the formulas is adjusted to the optimized transparency.
[0130] Back to Figure 1 In step 105, the image processing model may be trained by adjusting parameters of the image processing model based on the loss.
[0131] The following is a specific example to illustrate the training method of the image processing model disclosed in the present invention. Figure 2 It is a framework diagram of an image processing model according to an exemplary embodiment.
[0132] refer to Figure 2 First, a training image can be obtained and input into the trained instance segmentation model to obtain the mask of each of the three instances (three characters, corresponding to instance 1, instance 2, and instance 3, respectively). Each instance is marked with true transparency.
[0133] Next, three masks of each instance (3 masks of instance 1, 3 masks of instance 2, and 3 masks of instance 3) can be generated based on the mask of each instance. Specifically, based on the mask of each instance, a target mask, a reference mask, and a background mask of each instance are generated. The target mask, the reference mask, and the background mask can be seen in Figure 2 Shown in the lower left (the white part indicates the mask area).
[0134] Then, for each instance, the training image and the three masks can be input into the image processing model to obtain its three transparencies. Specifically, the training image and the three masks of instance 1 can be input into the image processing model to obtain the three transparencies of instance 1, the training image and the three masks of instance 2 can be input into the image processing model to obtain the three transparencies of instance 2, and the training image and the three masks of instance 3 can be input into the image processing model to obtain the three transparencies of instance 3, where the three transparencies are target transparency, reference transparency, and background transparency. Target transparency, reference transparency, and background transparency can be seen Figure 2 As shown in the lower right (transparency refers to the white area). Here, the image processing model can simultaneously generate image features corresponding to the instance.
[0135] Next, instance optimization may be performed, specifically: for each instance, the image features corresponding to the current instance are optimized based on the image features corresponding to the remaining instances, to obtain the optimized image features corresponding to the current instance.
[0136] Finally, the optimized image features corresponding to each instance can be input into the image processing model, and the optimized target transparency, reference transparency and background transparency corresponding to each instance can be output. Based on the loss calculation process in the above step 104, the loss calculation is performed according to the target transparency, reference transparency, background transparency, and the optimized target transparency, reference transparency and background transparency, and the parameters of the image processing model are adjusted to complete the training of the image processing model.
[0137] According to the exemplary embodiments of the present disclosure, quantitative and qualitative tests can be performed on the image processing model. First, for quantitative testing, the RWP636 dataset can be used. The compared algorithms include existing instance segmentation models (MaskRCNN and CascadePSP) and existing cutout models (GCA, SIM, FBA, MaskGuided). The comparative test found that the image processing model in the exemplary embodiments of the present disclosure achieved optimal performance under multiple measurement indicators on the dataset. In addition, for qualitative testing, Figure 3 FIG. 1 is a schematic diagram showing a qualitative test comparison according to an exemplary embodiment. Figure 3, the horizontal rows from top to bottom are the processing processes of 5 different images, and the vertical rows from left to right represent the original image, the processing results of MaskRCNN, the processing results of CascadePSP, the processing results of SIM, the processing results of MaskGuided, the processing results of the image processing model in the exemplary embodiment of the present disclosure, and the local enlarged image of the original image, the local processing results of CascadePSP, the local processing results of MaskGuided, and the local processing results of the image processing model in the exemplary embodiment of the present disclosure. Figure 3 It can be seen that for very difficult situations, such as hair mixed between two instances, or multiple instances occluding each other, the image processing model in the exemplary embodiments of the present disclosure can predict accurate results.
[0138] Figure 4 This is a flowchart of an image processing method according to an exemplary embodiment. The exemplary embodiment of the present disclosure can perform image processing based on an image processing model, obtain the mask of each instance, and obtain the transparency of each instance based on the image processing model, so that instance segmentation and instance extraction can be combined, which can solve the problem in the prior art that when there is occlusion between instances or they are close together, it is impossible to clearly distinguish each instance, and complete a fine instance cutout. Figure 4 In step 401, an image to be processed may be obtained, wherein the image to be processed includes at least one instance.
[0139] According to an exemplary embodiment of the present disclosure, an example of an image to be processed may be an individual person or an individual object for extraction and subsequent editing processing (ie, cutout).
[0140] At step 402 , a mask for each instance of at least one instance may be obtained.
[0141] According to an exemplary embodiment of the present disclosure, the mask of each instance can be obtained through instance segmentation to perform instance segmentation, and the existing deep learning method can be selected to implement it. For example, the image to be processed can be input into a trained instance segmentation model to obtain the mask of each instance of at least one instance. Here, the trained instance segmentation model can be an existing structure such as MaskRCNN, or other models that can perform instance segmentation, and the present disclosure does not limit this.
[0142] In step 403, for each instance, based on the mask of at least one instance, a target mask, a reference mask and a background mask of the current instance may be generated, and the image to be processed, the target mask, the reference mask and the background mask may be input into the image processing model to obtain the target transparency, the reference transparency and the background transparency of the current instance, wherein the target mask represents the mask of the current instance, the reference mask represents the masks of the remaining instances except the current instance, the background mask represents the mask of the remaining parts of the image to be processed except at least one instance, the target transparency is the transparency corresponding to the target mask, the reference transparency is the transparency corresponding to the reference mask, and the background transparency is the transparency corresponding to the background mask.
[0143] It should be noted that the image processing model can be composed of Figure 1 The training method of the image processing model shown is trained.
[0144] According to an exemplary embodiment of the present disclosure, a corresponding relationship between each instance and the image to be processed can be established. Then, from the perspective of the instance, the image to be processed can be divided into the cost instance, the remaining instances other than the instance, and the image portion other than the instance. In other words, the entire image to be processed can be divided into the target instance (current instance), the reference instance (the remaining instances other than the current instance), and the background (the remaining portion of the image to be processed other than the instance). Based on the above analysis, an exemplary embodiment of the present disclosure can generate a tri-mask of the current instance based on the mask of at least one instance, and the tri-mask includes a target mask, a reference mask, and a background mask.
[0145] For example, for the i-th instance of the image to be processed, three masks can be generated by the above equations (1) to (3).
[0146] Based on the above description, the above three masks and the image to be processed can be input into the image processing model to predict the transparency and obtain tri-matte, which includes target transparency, reference transparency and background transparency, corresponding to the target mask, reference mask and background mask respectively. The image processing model in the exemplary embodiment of the present disclosure can be, but not limited to, U-Net, ResNet34, ResNet50, ResNet101, MobileNet series. For example, when higher image processing accuracy is required, ResNet50 or ResNet101 can be used, and when faster processing speed is required, MobileNet series can be used. In the case where the image processing model is U-Net, it includes an encoder for encoding input items into features of different scales, and then based on a decoder, it reconstructs the original scale layer by layer from small to large according to these features, and finally sends the reconstructed features to a transparency prediction head composed of convolutional layers for prediction.
[0147] The image processing model in the exemplary embodiment of the present disclosure may include a feature extraction module and a transparency prediction module. Then, with respect to the structure of the image processing model, when predicting each transparency, the prediction may be carried out in a division of labor and cooperation among the modules. Specifically, the following may be done: first, the image to be processed, the target mask, the reference mask and the background mask are input into the feature extraction module to obtain the image features corresponding to the current instance. Then, the image features are input into the transparency prediction module to obtain the target transparency, the reference transparency and the background transparency of the current instance. It should be noted here that the image features may be the features generated by the last layer of the image processing model.
[0148] In step 404 , each instance may be extracted from the image to be processed based on the target transparency, the reference transparency, and the background transparency of each instance.
[0149] According to an exemplary embodiment of the present disclosure, an existing method may be used to extract instances, such as multiplying the image to be processed and the corresponding transparency to extract each instance, thereby completing the cutout operation.
[0150] For the above prediction process, although the guidance information of the remaining instances and the background has been introduced by inputting the reference mask and the background mask, there is no information interaction between the instances during the prediction process, which may cause the tri-matte predicted by each instance of the same image to be processed to be inconsistent. For example, each instance of the same image to be processed will predict the background transparency. In theory, the background transparency predicted by each instance should be consistent, but due to the difference in input information, the obtained background transparency may be inconsistent. In order to solve the problem of inconsistent predictions caused by information differences, after the transparency prediction module predicts each transparency, instance optimization can be performed to allow the maximum degree of freedom of information interaction between instances, and the transparency is predicted again based on the result after the interaction to obtain the optimized transparency. This process is described in detail below.
[0151] Specifically, for each instance, the target transparency, reference transparency, and background transparency of the remaining instances can be obtained, and the target transparency, reference transparency, and background transparency of the current instance can be optimized based on the target transparency, reference transparency, and background transparency of the remaining instances to obtain the optimized target transparency, reference transparency, and background transparency of the current instance. According to an embodiment, by optimizing the target transparency, reference transparency, and background transparency of the current instance, and optimizing based on the target transparency, reference transparency, and background transparency of the remaining instances, information interaction can be performed between instances.
[0152] Furthermore, the image features obtained in the feature extraction module can be used to implement the optimization. Then, for each instance, the following operations can be performed: first, the image features corresponding to at least one instance are obtained. Then, the image features corresponding to the current instance are optimized based on the image features corresponding to the remaining instances to obtain the optimized image features corresponding to the current instance. Finally, the optimized image features corresponding to the current instance are input into the transparency prediction module to obtain the optimized target transparency, reference transparency and background transparency corresponding to the current instance. According to the embodiment, the target transparency, reference transparency and background transparency of the current instance can be optimized based on the image features obtained by the feature extraction module.
[0153] Here, the image features corresponding to the current instance can be optimized in the order of feature decomposition, feature propagation, and feature aggregation. For feature propagation, the decomposed features of each instance can be propagated to the remaining instances, so that each instance receives the decomposed features from all the remaining instances. For each image to be processed, the number of instances corresponding to it may be different, so the embodiment of the present disclosure adopts a fusion method to perform feature propagation.
[0154] First, for feature decomposition, image features corresponding to each instance of at least one instance may be decomposed into target sub-features, reference sub-features, and background sub-features.
[0155] Here, the image feature corresponding to each instance can be combined with the target transparency of each instance to obtain the target sub-feature of each instance. The image feature corresponding to each instance and the reference transparency of each instance can be combined to obtain the reference sub-feature of each instance. The image feature corresponding to each instance and the background transparency of each instance can be combined to obtain the background sub-feature of each instance.
[0156] For example, the image feature corresponding to each instance may be multiplied by the target transparency, reference transparency and background transparency corresponding to the instance respectively to obtain the target sub-feature, reference sub-feature and background sub-feature corresponding to the corresponding instance.
[0157] Next, for feature propagation, the target sub-features, reference sub-features, and background sub-features corresponding to the current instance may be feature fused based on the target sub-features, reference sub-features, and background sub-features corresponding to the remaining instances to obtain optimized target sub-features, reference sub-features, and background sub-features corresponding to the current instance. For example, an exemplary embodiment of the present disclosure may perform feature fusion in a feature reduction manner.
[0158] Here, based on the target sub-features and reference sub-features corresponding to each instance, feature reduction may be performed on the target sub-features corresponding to the current instance to obtain an optimized target sub-feature corresponding to the current instance.
[0159] For example, the optimized target sub-feature corresponding to the i-th instance can be obtained through the above formula (6).
[0160] Based on the reference sub-feature corresponding to the current instance and the target sub-features corresponding to the remaining instances, feature reduction may be performed on the reference sub-feature corresponding to the current instance to obtain an optimized reference sub-feature corresponding to the current instance.
[0161] For example, the optimized reference sub-feature corresponding to the i-th instance can be obtained through the above formula (7).
[0162] Based on the background sub-features corresponding to each instance, feature reduction may be performed on the background sub-features corresponding to the current instance to obtain optimized background sub-features corresponding to the current instance.
[0163] For example, the optimized background sub-feature corresponding to the i-th instance can be obtained through the above formula (8).
[0164] According to the embodiment, the optimized target sub-feature, optimized reference sub-feature and optimized background sub-feature corresponding to the current instance can be obtained by feature reduction, and feature fusion can be achieved when the number of instances of the image input to the image processing model is different.
[0165] Finally, for feature aggregation, the optimized target sub-features, reference sub-features, and background sub-features corresponding to the current instance may be aggregated to obtain the optimized image features corresponding to the current instance. It should be noted that in the embodiments of the present disclosure, the aggregation operation may be implemented through a convolutional layer.
[0166] According to the embodiment, through the process of feature decomposition, feature propagation and feature aggregation, it is possible to optimize the target transparency, reference transparency and background transparency of the current instance based on the image features.
[0167] Based on the above optimization of transparency, each instance can be extracted from the image to be processed based on the optimized target transparency, reference transparency and background transparency of each instance to complete fine instance cutout.
[0168] Figure 5 is a block diagram of a training device for an image processing model according to an exemplary embodiment. Figure 5 The training device 500 includes a first image acquisition unit 501, a first mask acquisition unit 502, a first model prediction unit 503, a loss calculation unit 504 and a parameter adjustment unit 505.
[0169] The first image acquisition unit 501 can acquire a training image, wherein the training image includes at least one instance, and each instance is marked with true transparency.
[0170] According to an exemplary embodiment of the present disclosure, in the process of training the image processing model, the training images used may include many, Figure 5 The training apparatus shown in the figure is described with only one training image. For all images used to train the image processing model, the following can be used: Figure 5 The device configuration shown is not described in detail here. The training image can be, but is not limited to, an image containing at least one instance, which can be an individual person or an individual object for extraction and subsequent editing (i.e., cutout). For the exemplary embodiments of the present disclosure, for each instance of each training image, a real transparency is marked, and the real transparency can be used as an accurate real value for training to adjust the parameters of the model in combination with the predicted value of the model to complete the training of the model.
[0171] The training device 500 may first perform instance segmentation on at least one instance in the training image, and then perform subsequent transparency prediction on each instance. The following describes instance segmentation and transparency prediction respectively through the first mask acquisition unit 502 and the first model prediction unit 503.
[0172] The first mask acquisition unit 502 may acquire a mask of each instance of at least one instance.
[0173] According to an exemplary embodiment of the present disclosure, the first mask acquisition unit 502 can obtain the mask of each instance through instance segmentation to perform instance segmentation, which can be implemented by existing deep learning methods. For example, the first mask acquisition unit 502 can input the training image into a trained instance segmentation model to obtain the mask of each instance. Here, the trained instance segmentation model can be an existing structure such as MaskRCNN, or other models that can perform instance segmentation, and the present disclosure does not limit this.
[0174] The first model prediction unit 503 can generate a target mask, a reference mask and a background mask of the current instance for each instance based on the mask of at least one instance, and input the training image, the target mask, the reference mask and the background mask into the image processing model to obtain the target transparency, the reference transparency and the background transparency of the current instance, wherein the target mask represents the mask of the current instance, the reference mask represents the masks of the remaining instances except the current instance, the background mask represents the mask of the remaining parts except at least one instance in the training image, the target transparency is the transparency corresponding to the target mask, the reference transparency is the transparency corresponding to the reference mask, and the background transparency is the transparency corresponding to the background mask.
[0175] According to an exemplary embodiment of the present disclosure, a correspondence between each instance and a training image can be established. Then, from the perspective of the instance, the training image can be distinguished into the cost instance, the remaining instances other than the instance, and the image portion other than the instance. That is, the entire training image can be distinguished into a target instance (current instance), a reference instance (the remaining instances other than the current instance), and a background (the remaining portion of the training image other than the instance). Based on the above analysis, the first model prediction unit 503 can generate a tri-mask of the current instance based on the mask of at least one instance, and the tri-mask includes a target mask, a reference mask, and a background mask.
[0176] For example, for the i-th instance of the training image, three masks can be generated by the above equations (1) to (3).
[0177] Based on the above description, the first model prediction unit 503 can input the above three masks and training images into the image processing model to predict transparency and obtain tri-matte, which includes target transparency, reference transparency and background transparency, corresponding to the target mask, reference mask and background mask respectively. The image processing model in the exemplary embodiment of the present disclosure can be, but not limited to, U-Net, ResNet34, ResNet50, ResNet101, MobileNet series. For example, when higher image processing accuracy is required, ResNet50 or ResNet101 can be used, and when faster processing speed is required, MobileNet series can be used. In the case where the image processing model is U-Net, it includes an encoder for encoding input items into features of different scales, and then based on a decoder, it reconstructs the features layer by layer from small to large to the original scale, and finally sends the reconstructed features to a transparency prediction head composed of convolutional layers for prediction.
[0178] The image processing model in the exemplary embodiment of the present disclosure may include a feature extraction module and a transparency prediction module. Then, with respect to the structure of the image processing model, when predicting each transparency, the prediction may be carried out in a division of labor and cooperation among the modules. Specifically, the first model prediction unit 503 first inputs the training image, target mask, reference mask and background mask into the feature extraction module to obtain the image features corresponding to the current instance. Then the first model prediction unit 503 inputs the image features into the transparency prediction module to obtain the target transparency, reference transparency and background transparency of the current instance. It should be noted here that the image features may be the features generated by the last layer of the image processing model.
[0179] Back to Figure 5 The loss calculation unit 504 may calculate the loss based on the target transparency, reference transparency, background transparency and real transparency of each instance.
[0180] According to an exemplary embodiment of the present disclosure, for each instance of a training image, the loss may include four parts, namely, the first minimum absolute deviation, the first Laplace loss, the first synthesis constraint loss and the first transparency constraint loss. Then the loss may be the sum of the above four losses of all instances. The specific process of calculating the loss is specifically explained below.
[0181] First, the loss calculation unit 504 may determine the first minimum absolute deviation between the target transparency, reference transparency, background transparency and true transparency of each instance. It should be noted that the true transparency in the exemplary embodiment of the present disclosure may include true target transparency, true reference transparency and true background transparency.
[0182] Then, the loss calculation unit 504 may impose smoothness constraints on the parameters of the image processing model based on the target transparency, reference transparency, background transparency, and true transparency of each instance to obtain the first Laplace loss of each instance.
[0183] Next, the loss calculation unit 504 can obtain the first synthesis constraint loss of each instance based on the target synthesis part, reference synthesis part, background synthesis part and training image of each instance, wherein the target synthesis part is obtained based on the target transparency and the target instance foreground image, the reference synthesis part is obtained based on the reference transparency and the reference instance foreground image, the background synthesis part is obtained based on the background transparency and the background image, and the target instance foreground image, the reference instance foreground image and the background image are all determined based on the actual transparency and the training image.
[0184] For example, the first synthetic constraint loss can be obtained by the above formula (4).
[0185] Furthermore, the loss calculation unit 504 may obtain the first transparency constraint loss of each instance based on the target transparency, the reference transparency, and the background transparency of each instance.
[0186] For example, the first transparency constraint loss can be obtained by the above formula (5).
[0187] Finally, the loss calculation unit 504 may determine the loss based on the first minimum absolute deviation, the first Laplace loss, the first synthesis constraint loss, and the first transparency constraint loss of each instance. For example, but not limited to, the loss calculation unit 504 may add the first minimum absolute deviation, the first Laplace loss, the first synthesis constraint loss, and the first transparency constraint loss of each instance to obtain the loss.
[0188] It should be noted that the order in which the loss calculation unit 504 obtains the four losses of each instance is merely illustrative, and other orders may be used for calculation, and the present disclosure does not impose any limitation on this.
[0189] According to the embodiment, through the above-mentioned first synthesis constraint loss and first transparency constraint loss, the loss can be determined for the predicted target transparency, reference transparency and background transparency in a targeted manner.
[0190] For the above prediction process, although the guidance information of the remaining instances and the background has been introduced by inputting the reference mask and the background mask, there is no information interaction between the instances during the prediction process, which may cause the tri-matte predicted by each instance of the same training image to be inconsistent. For example, each instance of the same training image will predict the background transparency. In theory, the background transparency predicted by each instance should be consistent, but due to the difference in input information, the obtained background transparency may be inconsistent. In order to solve the problem of inconsistent predictions caused by information differences, after the transparency prediction module predicts each transparency, instance optimization (multi-instance refinement) can be performed to enable the maximum degree of freedom of information interaction between instances, and the transparency is predicted again based on the result after the interaction to obtain the optimized transparency, which is explained in detail below.
[0191] Specifically, the training device 500 may be configured to include a first optimization unit, which may obtain the target transparency, reference transparency, and background transparency of the remaining instances for each instance, and optimize the target transparency, reference transparency, and background transparency of the current instance based on the target transparency, reference transparency, and background transparency of the remaining instances to obtain the optimized target transparency, reference transparency, and background transparency of the current instance. According to an embodiment, by optimizing the target transparency, reference transparency, and background transparency of the current instance, and optimizing based on the target transparency, reference transparency, and background transparency of the remaining instances, information exchange between instances is enabled.
[0192] Furthermore, the first optimization unit may use the image features obtained in the feature extraction module to implement the optimization. Then, the first optimization unit may perform the following operations for each instance: first, obtain the image features corresponding to at least one instance. Then, based on the image features corresponding to the remaining instances, the image features corresponding to the current instance are optimized to obtain the optimized image features corresponding to the current instance. Finally, the optimized image features corresponding to the current instance are input into the transparency prediction module to obtain the optimized target transparency, reference transparency, and background transparency corresponding to the current instance. According to an embodiment, the target transparency, reference transparency, and background transparency of the current instance may be optimized based on the image features obtained by the feature extraction module.
[0193] Here, the first optimization unit can optimize the image features corresponding to the current instance in the order of feature decomposition, feature propagation and feature aggregation. For feature propagation, the decomposed features of each instance can be propagated to the remaining instances respectively, so that each instance receives the decomposed features from all the remaining instances. For each training image, the number of instances corresponding to it may be different, so the first optimization unit adopts a fusion method to perform feature propagation.
[0194] First, for feature decomposition, the first optimization unit may decompose the image features corresponding to each instance of the at least one instance into target sub-features, reference sub-features, and background sub-features.
[0195] Here, the first optimization unit may combine the image feature corresponding to each instance with the target transparency of each instance to obtain the target sub-feature of each instance. The first optimization unit may combine the image feature corresponding to each instance with the reference transparency of each instance to obtain the reference sub-feature of each instance. The first optimization unit may combine the image feature corresponding to each instance with the background transparency of each instance to obtain the background sub-feature of each instance.
[0196] For example, the first optimization unit may multiply the image feature corresponding to each instance by the target transparency, reference transparency and background transparency corresponding to the instance, respectively, to obtain the target sub-feature, reference sub-feature and background sub-feature corresponding to the corresponding instance.
[0197] Next, for feature propagation, the first optimization unit may perform feature fusion on the target sub-features, reference sub-features, and background sub-features corresponding to the current instance based on the target sub-features, reference sub-features, and background sub-features corresponding to the remaining instances, to obtain the optimized target sub-features, reference sub-features, and background sub-features corresponding to the current instance. For example, an exemplary embodiment of the present disclosure may perform feature fusion in a feature reduction manner.
[0198] Here, the first optimization unit may perform feature reduction on the target sub-feature corresponding to the current instance based on the target sub-feature and the reference sub-feature corresponding to each instance to obtain the optimized target sub-feature corresponding to the current instance.
[0199] For example, the optimized target sub-feature corresponding to the i-th instance can be obtained through the above formula (6).
[0200] The first optimization unit may perform feature reduction on the reference sub-feature corresponding to the current instance based on the reference sub-feature corresponding to the current instance and the target sub-features corresponding to the remaining instances to obtain an optimized reference sub-feature corresponding to the current instance.
[0201] For example, the optimized reference sub-feature corresponding to the i-th instance can be obtained through the above formula (7).
[0202] The first optimization unit may perform feature reduction on the background sub-features corresponding to the current instance based on the background sub-features corresponding to each instance to obtain optimized background sub-features corresponding to the current instance.
[0203] For example, the optimized background sub-feature corresponding to the i-th instance can be obtained through the above formula (8).
[0204] According to an embodiment, the first optimization unit can obtain the optimized target sub-features, optimized reference sub-features and optimized background sub-features corresponding to the current instance by feature reduction, and can realize feature fusion when the number of instances of the image input to the image processing model is different.
[0205] Finally, for feature aggregation, the first optimization unit may aggregate the optimized target sub-features, reference sub-features, and background sub-features corresponding to the current instance to obtain the optimized image features corresponding to the current instance. It should be noted that in the embodiments of the present disclosure, the aggregation operation may be implemented through a convolutional layer.
[0206] According to the embodiment, the first optimization unit can optimize the target transparency, reference transparency and background transparency of the current instance based on the image features through the processes of feature decomposition, feature propagation and feature aggregation.
[0207] Based on the above-mentioned optimization of transparency, the loss calculation unit 504 can first calculate the first loss based on the target transparency, reference transparency, background transparency and true transparency of each instance; then calculate the second loss based on the optimized target transparency, reference transparency, background transparency and true transparency of each instance; finally, determine the loss based on the first loss and the second loss.
[0208] According to an embodiment, the loss calculation unit 504 takes into account the target transparency, reference transparency, background transparency of each instance, as well as the optimized target transparency, reference transparency, and background transparency of each instance in the process of determining the loss, which can enable the trained image processing model to have more accurate prediction results.
[0209] Specifically, first, the loss calculation unit 504 can determine the first minimum absolute deviation between the target transparency, reference transparency, background transparency, and true transparency of each instance, and determine the second minimum absolute deviation between the optimized target transparency, reference transparency, background transparency, and true transparency of each instance.
[0210] Then, the loss calculation unit 504 may apply smoothness constraints on the parameters of the image processing model based on the target transparency, reference transparency, background transparency and true transparency of each instance to obtain the first Laplace loss of each instance, and apply smoothness constraints on the parameters of the image processing model based on the optimized target transparency, reference transparency, background transparency and true transparency of each instance to obtain the second Laplace loss of each instance.
[0211] Next, the loss calculation unit 504 can obtain the first synthesis constraint loss of each instance based on the target synthesis part, reference synthesis part, background synthesis part and training image of each instance, and obtain the second synthesis constraint loss of each instance based on the optimized target synthesis part, optimized reference synthesis part, optimized background synthesis part and training image of each instance, wherein the target synthesis part is obtained based on the target transparency and the target instance foreground image, the reference synthesis part is obtained based on the reference transparency and the reference instance foreground image, the background synthesis part is obtained based on the background transparency and the background image, the target instance foreground image, the reference instance foreground image and the background image are all determined based on the real transparency and the training image, the optimized target synthesis part is obtained based on the optimized target transparency and the target instance foreground image, the optimized reference synthesis part is obtained based on the optimized reference transparency and the reference instance foreground image, and the optimized background synthesis part is obtained based on the optimized background transparency and the background image.
[0212] Then, the loss calculation unit 504 can obtain the first transparency constraint loss of each instance based on the target transparency, reference transparency and background transparency of each instance, and obtain the second transparency constraint loss of each instance based on the optimized target transparency, reference transparency and background transparency of each instance.
[0213] Next, the loss calculation unit 504 may determine the first loss based on the first minimum absolute deviation, the first Laplace loss, the first synthetic constraint loss, and the first transparency constraint loss of each instance, and determine the second loss based on the second minimum absolute deviation, the second Laplace loss, the second synthetic constraint loss, and the second transparency constraint loss of each instance. For example, the loss calculation unit 504 may add the first minimum absolute deviation, the first Laplace loss, the first synthetic constraint loss, and the first transparency constraint loss of each instance to obtain the first loss; and may add the second minimum absolute deviation, the second Laplace loss, the second synthetic constraint loss, and the second transparency constraint loss of each instance to obtain the second loss.
[0214] Finally, the loss calculation unit 504 may determine the loss based on the first loss and the second loss. For example, the loss calculation unit 504 may add the first loss and the second loss to obtain the loss.
[0215] According to the embodiment, the loss calculation unit 504 can determine the loss of the predicted target transparency, reference transparency and background transparency, as well as the optimized target transparency, reference transparency and background transparency through the above-mentioned first synthetic constraint loss and first transparency constraint loss, second synthetic constraint loss and second transparency constraint loss.
[0216] It should be noted that the order in which the loss calculation unit 504 obtains the loss is only illustrative, and other orders may be used for calculation, and the present disclosure does not limit this. The second synthetic constraint loss and the second transparency constraint loss may respectively adopt the above formulas (4) and (5), and the corresponding transparency in the formulas is adjusted to the optimized transparency.
[0217] Back to Figure 5 The parameter adjustment unit 505 can train the image processing model by adjusting the parameters of the image processing model based on the loss.
[0218] Figure 6 The block diagram of an image processing device according to an exemplary embodiment is shown. The exemplary embodiment of the present disclosure can perform image processing based on an image processing model, obtain the mask of each instance, and obtain the transparency of each instance based on the image processing model, so that instance segmentation and instance extraction can be combined, which can solve the problem in the prior art that when there is occlusion between instances or they are close together, it is impossible to clearly distinguish each instance, and complete fine instance cutout. Figure 6 The image processing device 600 includes a second image acquisition unit 601, a second mask acquisition unit 602, a second model prediction unit 603 and an instance extraction unit 604.
[0219] The second image acquisition unit 601 can acquire an image to be processed, wherein the image to be processed includes at least one instance.
[0220] According to an exemplary embodiment of the present disclosure, an example of an image to be processed may be an individual person or an individual object for extraction and subsequent editing processing (ie, cutout).
[0221] The second mask acquisition unit 602 may acquire a mask of each instance of at least one instance.
[0222] According to an exemplary embodiment of the present disclosure, the second mask acquisition unit 602 can obtain the mask of each instance through instance segmentation to perform instance segmentation, and can be implemented through existing deep learning methods. For example, the second mask acquisition unit 602 can input the image to be processed into a trained instance segmentation model to obtain the mask of each instance of at least one instance. Here, the trained instance segmentation model can be an existing structure such as MaskRCNN, or other models that can perform instance segmentation, and the present disclosure does not limit this.
[0223] The second model prediction unit 603 can generate a target mask, a reference mask and a background mask of the current instance for each instance based on the mask of at least one instance, and input the image to be processed, the target mask, the reference mask and the background mask into the image processing model to obtain the target transparency, the reference transparency and the background transparency of the current instance, wherein the target mask represents the mask of the current instance, the reference mask represents the masks of the remaining instances except the current instance, the background mask represents the mask of the remaining parts of the image to be processed except at least one instance, the target transparency is the transparency corresponding to the target mask, the reference transparency is the transparency corresponding to the reference mask, and the background transparency is the transparency corresponding to the background mask.
[0224] It should be noted that the image processing model can be composed of Figure 1 The training method of the image processing model shown is trained.
[0225] According to an exemplary embodiment of the present disclosure, a corresponding relationship between each instance and the image to be processed can be established. Then, from the perspective of the instance, the image to be processed can be divided into the cost instance, the remaining instances other than the instance, and the image portion other than the instance. That is, the entire image to be processed can be divided into the target instance (current instance), the reference instance (the remaining instances other than the current instance), and the background (the remaining portion of the image to be processed other than the instance). Based on the above analysis, the second model prediction unit 603 can generate a tri-mask of the current instance based on the mask of at least one instance, and the tri-mask includes a target mask, a reference mask, and a background mask.
[0226] For example, for the i-th instance of the image to be processed, the second model prediction unit 603 may generate three masks through the above equations (1) to (3).
[0227] Based on the above description, the second model prediction unit 603 can input the above three masks and the image to be processed into the image processing model to predict the transparency and obtain tri-matte, which includes target transparency, reference transparency and background transparency, corresponding to the target mask, reference mask and background mask respectively. The image processing model in the exemplary embodiment of the present disclosure can be, but not limited to, U-Net, ResNet34, ResNet50, ResNet101, MobileNet series. For example, when higher image processing accuracy is required, ResNet50 or ResNet101 can be used, and when faster processing speed is required, MobileNet series can be used. In the case where the image processing model is U-Net, it includes an encoder for encoding input items into features of different scales, and then based on a decoder, it reconstructs the features layer by layer from small to large to the original scale, and finally sends the reconstructed features to a transparency prediction head composed of convolutional layers for prediction.
[0228] The image processing model in the exemplary embodiment of the present disclosure may include a feature extraction module and a transparency prediction module. Then, with respect to the structure of the image processing model, when predicting each transparency, the prediction may be carried out in a division of labor and cooperation among the modules. Specifically, the second model prediction unit 603 first inputs the image to be processed, the target mask, the reference mask and the background mask into the feature extraction module to obtain the image features corresponding to the current instance. Then the second model prediction unit 603 inputs the image features into the transparency prediction module to obtain the target transparency, the reference transparency and the background transparency of the current instance. It should be noted here that the image features may be the features generated by the last layer of the image processing model.
[0229] The instance extraction unit 604 may extract each instance from the image to be processed based on the target transparency, reference transparency and background transparency of each instance.
[0230] According to an exemplary embodiment of the present disclosure, the instance extraction unit 604 may extract instances using existing methods, such as multiplying the image to be processed and the corresponding transparency to extract each instance, thereby completing the cutout operation.
[0231] For the above prediction process, although the guidance information of the remaining instances and the background has been introduced by inputting the reference mask and the background mask, there is no information interaction between the instances during the prediction process, which may cause the tri-matte predicted by each instance of the same image to be processed to be inconsistent. For example, each instance of the same image to be processed will predict the background transparency. In theory, the background transparency predicted by each instance should be consistent, but due to the difference in input information, the obtained background transparency may be inconsistent. In order to solve the problem of inconsistent predictions caused by information differences, after the transparency prediction module predicts each transparency, instance optimization can be performed to allow the maximum degree of freedom of information interaction between instances, and the transparency is predicted again based on the result after the interaction to obtain the optimized transparency, which is explained in detail below.
[0232] Specifically, the image processing device 600 may be configured with a second optimization unit, and the second optimization unit may obtain the target transparency, reference transparency, and background transparency of the remaining instances for each instance, and optimize the target transparency, reference transparency, and background transparency of the current instance based on the target transparency, reference transparency, and background transparency of the remaining instances to obtain the optimized target transparency, reference transparency, and background transparency of the current instance. According to an embodiment, the second optimization unit optimizes the target transparency, reference transparency, and background transparency of the current instance, and optimizes the target transparency, reference transparency, and background transparency based on the target transparency, reference transparency, and background transparency of the remaining instances, so that information exchange can be performed between instances.
[0233] Furthermore, the second optimization unit may utilize the image features obtained in the feature extraction module to implement the optimization. Then, the second optimization unit may perform the following operations for each instance: First, obtain the image features corresponding to at least one instance. Then, optimize the image features corresponding to the current instance based on the image features corresponding to the remaining instances to obtain the optimized image features corresponding to the current instance. Finally, input the optimized image features corresponding to the current instance into the transparency prediction module to obtain the optimized target transparency, reference transparency, and background transparency corresponding to the current instance. According to an embodiment, the second optimization unit may optimize the target transparency, reference transparency, and background transparency of the current instance based on the image features obtained by the feature extraction module.
[0234] Here, the second optimization unit can optimize the image features corresponding to the current instance in the order of feature decomposition, feature propagation and feature aggregation. For feature propagation, the decomposed features of each instance can be propagated to the remaining instances respectively, so that each instance receives the decomposed features from all the remaining instances. For each image to be processed, the number of instances corresponding to it may be different, so the second optimization unit adopts a fusion method to perform feature propagation.
[0235] First, for feature decomposition, the second optimization unit may decompose the image features corresponding to each instance of the at least one instance into target sub-features, reference sub-features, and background sub-features.
[0236] Here, the second optimization unit may combine the image feature corresponding to each instance with the target transparency of each instance to obtain the target sub-feature of each instance. The second optimization unit may combine the image feature corresponding to each instance with the reference transparency of each instance to obtain the reference sub-feature of each instance. The second optimization unit may combine the image feature corresponding to each instance with the background transparency of each instance to obtain the background sub-feature of each instance.
[0237] For example, the second optimization unit may multiply the image feature corresponding to each instance by the target transparency, reference transparency and background transparency corresponding to the instance, respectively, to obtain the target sub-feature, reference sub-feature and background sub-feature corresponding to the corresponding instance.
[0238] Next, for feature propagation, the second optimization unit may perform feature fusion on the target sub-features, reference sub-features, and background sub-features corresponding to the current instance based on the target sub-features, reference sub-features, and background sub-features corresponding to the remaining instances, to obtain the optimized target sub-features, reference sub-features, and background sub-features corresponding to the current instance. For example, an exemplary embodiment of the present disclosure may perform feature fusion in a feature reduction manner.
[0239] Here, the second optimization unit may perform feature reduction on the target sub-feature corresponding to the current instance based on the target sub-feature and the reference sub-feature corresponding to each instance to obtain the optimized target sub-feature corresponding to the current instance.
[0240] For example, the second optimization unit can obtain the optimized target sub-feature corresponding to the i-th instance through the above formula (6).
[0241] The second optimization unit may perform feature reduction on the reference sub-feature corresponding to the current instance based on the reference sub-feature corresponding to the current instance and the target sub-features corresponding to the remaining instances to obtain an optimized reference sub-feature corresponding to the current instance.
[0242] For example, the second optimization unit can obtain the optimized reference sub-feature corresponding to the i-th instance through the above formula (7).
[0243] The second optimization unit may perform feature reduction on the background sub-features corresponding to the current instance based on the background sub-features corresponding to each instance to obtain optimized background sub-features corresponding to the current instance.
[0244] For example, the second optimization unit can obtain the optimized background sub-feature corresponding to the i-th instance through the above formula (8).
[0245] According to an embodiment, the second optimization unit can obtain the optimized target sub-features, optimized reference sub-features and optimized background sub-features corresponding to the current instance by feature reduction, and can realize feature fusion when the number of instances of the image input to the image processing model is different.
[0246] Finally, for feature aggregation, the second optimization unit may aggregate the optimized target sub-features, reference sub-features, and background sub-features corresponding to the current instance to obtain the optimized image features corresponding to the current instance. It should be noted that in the embodiments of the present disclosure, the aggregation operation may be implemented through a convolutional layer.
[0247] According to the embodiment, the second optimization unit can optimize the target transparency, reference transparency and background transparency of the current instance based on the image features through the processes of feature decomposition, feature propagation and feature aggregation.
[0248] Based on the above optimization of transparency, the instance extraction unit 604 may extract each instance from the image to be processed based on the optimized target transparency, reference transparency and background transparency of each instance to complete fine instance cutout.
[0249] Figure 7 is a block diagram of an electronic device 700 according to an exemplary embodiment.
[0250] Reference Figure 7 The electronic device 700 includes at least one memory 701 and at least one processor 702, wherein the at least one memory 701 stores a set of computer executable instructions. When the computer executable instruction set is executed by the at least one processor 702, the training method of the image processing model or the image processing method according to the present disclosure is executed.
[0251] As an example, the electronic device 700 may be a PC, a tablet device, a personal digital assistant, a smart phone, or other device capable of executing the above instruction set. Here, the electronic device 700 is not necessarily a single electronic device, but may also be any device or circuit capable of executing the above instruction (or instruction set) individually or in combination. The electronic device 700 may also be part of an integrated control system or system manager, or may be configured as a portable electronic device interconnected with a local or remote (e.g., via wireless transmission) interface.
[0252] In the electronic device 700, the processor 702 may include a central processing unit (CPU), a graphics processing unit (GPU), a programmable logic device, a dedicated processor system, a microcontroller or a microprocessor. As an example and not limitation, the processor may also include an analog processor, a digital processor, a microprocessor, a multi-core processor, a processor array, a network processor, etc.
[0253] The processor 702 may execute instructions or codes stored in the memory 701, which may also store data. Instructions and data may also be sent and received over a network via a network interface device, which may employ any known transmission protocol.
[0254] The memory 701 may be integrated with the processor 702, for example, by placing RAM or flash memory within an integrated circuit microprocessor or the like. In addition, the memory 701 may include a separate device, such as an external disk drive, a storage array, or any other storage device that can be used by a database system. The memory 701 and the processor 702 may be operatively coupled, or may communicate with each other, such as through an I / O port, a network connection, etc., so that the processor 702 can read files stored in the memory.
[0255] In addition, the electronic device 700 may further include a video display (such as a liquid crystal display) and a user interaction interface (such as a keyboard, a mouse, a touch input device, etc.) All components of the electronic device 700 may be connected to each other via a bus and / or a network.
[0256] According to an exemplary embodiment of the present disclosure, a computer-readable storage medium storing instructions may also be provided, wherein when the instructions are executed by at least one processor, the at least one processor is prompted to execute the training method of the image processing model or the image processing method according to the present disclosure. Examples of computer-readable storage media here include: read-only memory (ROM), random access programmable read-only memory (PROM), electrically erasable programmable read-only memory (EEPROM), random access memory (RAM), dynamic random access memory (DRAM), static random access memory (SRAM), flash memory, non-volatile memory, CD-ROM, CD-R, CD+R, CD-RW, CD+RW, DVD-ROM, DVD-R, DVD+R, DVD-RW, DVD+RW, DVD-RAM, BD-ROM, BD-R, BD-R LTH, BD-RE, Blu-ray or optical disk storage, hard disk drive (HDD), solid state drive (SSD), card storage (such as, multimedia card, secure digital (SD) card or extreme digital (XD) card), magnetic tape, floppy disk, magneto-optical data storage device, optical data storage device, hard disk, solid state disk and any other device, any other device is configured to store computer programs and any associated data, data files and data structures in a non-transitory manner and provide the computer programs and any associated data, data files and data structures to a processor or computer so that the processor or computer can execute the computer program. The computer program in the above-mentioned computer-readable storage medium can be run in an environment deployed in a computer device such as a client, a host, an agent device, a server, etc. In addition, in one example, the computer program and any associated data, data files and data structures are distributed on a networked computer system, so that the computer program and any associated data, data files and data structures are stored, accessed and executed in a distributed manner by one or more processors or computers.
[0257] According to an exemplary embodiment of the present disclosure, a computer program product may also be provided, and instructions in the computer program product may be executed by a processor of a computer device to complete the training method of the image processing model or the image processing method according to the present disclosure.
[0258] According to the training method and device of the image processing model and the image processing method and device disclosed in the present invention, by obtaining the mask of each instance and obtaining the transparency of each instance based on the image processing model, instance segmentation and instance extraction can be combined. By generating the target mask, reference mask and background mask of the instance, the image processing model can predict the target transparency, reference transparency and background transparency, so that the image processing model can not only learn the difference between the instance and the background, but also learn the difference between the instances, so as to obtain more accurate prediction results, and can solve the problem in the prior art that when the instances are blocked or close together, the instances cannot be clearly distinguished, and complete the fine instance cutout.
[0259] In addition, according to the training method and device of the image processing model and the image processing method and device disclosed in the present invention, instances can be optimized, and information interaction can be performed between instances to obtain more accurate and precise prediction results.
[0260] In addition, the training method and device of the image processing model and the image processing method and device according to the present invention can be easily applied to many practical scenarios, such as extracting only the image of the speaker in a meeting scene and ignoring other people in the background, or realizing personalized editing for individuals in picture editing applications.
[0261] Those skilled in the art will readily appreciate other embodiments of the present disclosure after considering the specification and practicing the invention disclosed herein. This application is intended to cover any variations, uses, or adaptations of the present disclosure that follow the general principles of the present disclosure and include common knowledge or customary techniques in the art that are not disclosed in the present disclosure. The specification and examples are intended to be exemplary only, and the true scope and spirit of the present disclosure are indicated by the following claims.
[0262] It should be understood that the present disclosure is not limited to the exact structures that have been described above and shown in the drawings, and that various modifications and changes may be made without departing from the scope thereof. The scope of the present disclosure is limited only by the appended claims.
Claims
1. A training method for an image processing model, characterized in that: include: Acquire a training image, wherein the training image includes at least one instance, and each instance is marked with a true transparency; obtaining a mask for each instance of the at least one instance; For each instance, based on the mask of the at least one instance, a target mask, a reference mask and a background mask of the current instance are generated, and the training image, the target mask, the reference mask and the background mask are input into an image processing model to obtain a target transparency, a reference transparency and a background transparency of the current instance, wherein the target mask represents the mask of the current instance, the reference mask represents the masks of the remaining instances except the current instance, the background mask represents the mask of the remaining parts of the training image except the at least one instance, the target transparency is the transparency corresponding to the target mask, the reference transparency is the transparency corresponding to the reference mask, and the background transparency is the transparency corresponding to the background mask; Calculating a loss based on the target transparency, the reference transparency, the background transparency, and the true transparency of each instance; The image processing model is trained by adjusting parameters of the image processing model based on the loss.
2. The training method according to claim 1, characterized in that: The image processing model includes a feature extraction module and a transparency prediction module; The step of inputting the training image, the target mask, the reference mask and the background mask into an image processing model to obtain the target transparency, the reference transparency and the background transparency of the current instance includes: Inputting the training image, the target mask, the reference mask and the background mask into the feature extraction module to obtain image features corresponding to the current instance; The image features are input into the transparency prediction module to obtain the target transparency, reference transparency and background transparency of the current instance.
3. The training method according to claim 1 or 2, characterized in that: Also includes: For each instance, the target transparency, reference transparency and background transparency of the remaining instances are obtained, and the target transparency, reference transparency and background transparency of the current instance are optimized based on the target transparency, reference transparency and background transparency of the remaining instances to obtain the optimized target transparency, reference transparency and background transparency of the current instance.
4. The training method according to claim 3, characterized in that: The calculating of the loss based on the target transparency, the reference transparency, the background transparency and the real transparency of each instance includes: Calculating a first loss based on the target transparency, the reference transparency, the background transparency, and the true transparency of each instance; Calculating a second loss based on the optimized target transparency, the reference transparency, the background transparency, and the true transparency of each instance; The loss is determined based on the first loss and the second loss.
5. The training method according to claim 3, characterized in that: The optimizing the target transparency, reference transparency and background transparency of the current instance based on the target transparency, reference transparency and background transparency of the remaining instances to obtain the optimized target transparency, reference transparency and background transparency of the current instance includes: For each instance, do the following: Acquire an image feature corresponding to the at least one instance; Optimizing the image features corresponding to the current instance based on the image features corresponding to the remaining instances to obtain optimized image features corresponding to the current instance; The optimized image features corresponding to the current instance are input into the transparency prediction module to obtain the optimized target transparency, reference transparency and background transparency corresponding to the current instance.
6. The training method according to claim 5, characterized in that: The optimizing the image features corresponding to the current instance based on the image features corresponding to the remaining instances to obtain the optimized image features corresponding to the current instance includes: Decomposing the image features corresponding to each instance of the at least one instance into target sub-features, reference sub-features and background sub-features; Based on the target sub-features, reference sub-features and background sub-features corresponding to the remaining instances, feature fusion is performed on the target sub-features, reference sub-features and background sub-features corresponding to the current instance to obtain optimized target sub-features, reference sub-features and background sub-features corresponding to the current instance; The optimized target sub-feature, reference sub-feature and background sub-feature corresponding to the current instance are aggregated to obtain the optimized image feature corresponding to the current instance.
7. The training method according to claim 6, characterized in that: The step of decomposing the image features corresponding to each instance of the at least one instance into target sub-features, reference sub-features and background sub-features comprises: Combining the image feature corresponding to each instance with the target transparency of each instance to obtain the target sub-feature of each instance; Combining the image feature corresponding to each instance and the reference transparency of each instance to obtain the reference sub-feature of each instance; The image feature corresponding to each instance and the background transparency of each instance are combined to obtain the background sub-feature of each instance.
8. The training method according to claim 6, characterized in that: The performing feature fusion on the target sub-feature, the reference sub-feature and the background sub-feature corresponding to the current instance based on the target sub-feature, the reference sub-feature and the background sub-feature corresponding to the remaining instances includes: Based on the target sub-features and reference sub-features corresponding to each instance, feature reduction is performed on the target sub-feature corresponding to the current instance to obtain an optimized target sub-feature corresponding to the current instance; Based on the reference sub-feature corresponding to the current instance and the target sub-features corresponding to the remaining instances, feature reduction is performed on the reference sub-feature corresponding to the current instance to obtain an optimized reference sub-feature corresponding to the current instance; Based on the background sub-features corresponding to each instance, feature reduction is performed on the background sub-features corresponding to the current instance to obtain optimized background sub-features corresponding to the current instance.
9. The training method according to claim 1, characterized in that: The calculating of the loss based on the target transparency, the reference transparency, the background transparency and the real transparency of each instance includes: determining a first minimum absolute deviation among the target transparency, the reference transparency, the background transparency, and the true transparency of each instance; Based on the target transparency, the reference transparency, the background transparency and the true transparency of each instance, smoothness constraints are imposed on the parameters of the image processing model to obtain the first Laplace loss of each instance; Based on the target synthesis part, the reference synthesis part, the background synthesis part and the training image of each instance, a first synthesis constraint loss of each instance is obtained, wherein the target synthesis part is obtained based on the target transparency and the target instance foreground image, the reference synthesis part is obtained based on the reference transparency and the reference instance foreground image, the background synthesis part is obtained based on the background transparency and the background image, and the target instance foreground image, the reference instance foreground image and the background image are all determined based on the real transparency and the training image; Based on the target transparency, the reference transparency and the background transparency of each instance, obtaining a first transparency constraint loss of each instance; The loss is determined based on the first minimum absolute deviation of each instance, the first Laplace loss, the first synthesis constraint loss, and the first transparency constraint loss.
10. The training method according to claim 4, characterized in that: The calculating of the loss based on the target transparency, the reference transparency, the background transparency and the real transparency of each instance includes: Determining a first minimum absolute deviation among the target transparency, the reference transparency, the background transparency, and the true transparency of each instance, and determining a second minimum absolute deviation among the optimized target transparency, the reference transparency, the background transparency, and the true transparency of each instance; Based on the target transparency, the reference transparency, the background transparency and the true transparency of each instance, smooth constraints are imposed on the parameters of the image processing model to obtain a first Laplace loss of each instance, and based on the optimized target transparency, the reference transparency, the background transparency and the true transparency of each instance, smooth constraints are imposed on the parameters of the image processing model to obtain a second Laplace loss of each instance; Based on the target synthesis part, the reference synthesis part, the background synthesis part and the training image of each instance, a first synthesis constraint loss of each instance is obtained, and based on the optimized target synthesis part, the optimized reference synthesis part, the optimized background synthesis part and the training image of each instance, a second synthesis constraint loss of each instance is obtained, wherein the target synthesis part is obtained based on the target transparency and the target instance foreground image, the reference synthesis part is obtained based on the reference transparency and the reference instance foreground image, the background synthesis part is obtained based on the background transparency and the background image, the target instance foreground image, the reference instance foreground image and the background image are all determined based on the real transparency and the training image, the optimized target synthesis part is obtained based on the optimized target transparency and the target instance foreground image, the optimized reference synthesis part is obtained based on the optimized reference transparency and the reference instance foreground image, and the optimized background synthesis part is obtained based on the optimized background transparency and the background image; Based on the target transparency, the reference transparency and the background transparency of each instance, a first transparency constraint loss of each instance is obtained, and based on the optimized target transparency, the reference transparency and the background transparency of each instance, a second transparency constraint loss of each instance is obtained; determining the first loss based on the first minimum absolute deviation, the first Laplace loss, the first composite constraint loss, and the first transparency constraint loss of each of the instances, and determining the second loss based on the second minimum absolute deviation, the second Laplace loss, the second composite constraint loss, and the second transparency constraint loss of each of the instances; The loss is determined based on the first loss and the second loss.
11. An image processing method, characterized in that: include: Acquire an image to be processed, wherein the image to be processed includes at least one instance; obtaining a mask for each instance of the at least one instance; For each instance, based on the mask of the at least one instance, a target mask, a reference mask and a background mask of the current instance are generated, and the image to be processed, the target mask, the reference mask and the background mask are input into an image processing model to obtain a target transparency, a reference transparency and a background transparency of the current instance, wherein the target mask represents the mask of the current instance, the reference mask represents the masks of the remaining instances except the current instance, the background mask represents the mask of the remaining parts of the image to be processed except the at least one instance, the target transparency is the transparency corresponding to the target mask, the reference transparency is the transparency corresponding to the reference mask, and the background transparency is the transparency corresponding to the background mask; Each instance is extracted from the image to be processed based on the target transparency, the reference transparency and the background transparency of each instance.
12. The image processing method according to claim 11, characterized in that: The image processing model includes a feature extraction module and a transparency prediction module; The step of inputting the image to be processed, the target mask, the reference mask and the background mask into an image processing model to obtain the target transparency, the reference transparency and the background transparency of the current instance includes: Inputting the image to be processed, the target mask, the reference mask and the background mask into the feature extraction module to obtain image features corresponding to the current instance; The image features are input into the transparency prediction module to obtain the target transparency, reference transparency and background transparency of the current instance.
13. The image processing method according to claim 11 or 12, characterized in that: Also includes: For each instance, the target transparency, reference transparency and background transparency of the remaining instances are obtained, and the target transparency, reference transparency and background transparency of the current instance are optimized based on the target transparency, reference transparency and background transparency of the remaining instances to obtain the optimized target transparency, reference transparency and background transparency of the current instance.
14. The image processing method according to claim 13, characterized in that: The step of extracting each instance from the image to be processed based on the target transparency, the reference transparency, and the background transparency of each instance comprises: Each instance is extracted from the image to be processed based on the optimized target transparency, reference transparency and background transparency of each instance.
15. The image processing method according to claim 13, characterized in that: The optimizing the target transparency, reference transparency and background transparency of the current instance based on the target transparency, reference transparency and background transparency of the remaining instances to obtain the optimized target transparency, reference transparency and background transparency of the current instance includes: For each instance, do the following: Acquire an image feature corresponding to the at least one instance; Optimizing the image features corresponding to the current instance based on the image features corresponding to the remaining instances to obtain optimized image features corresponding to the current instance; The optimized image features corresponding to the current instance are input into the transparency prediction module to obtain the optimized target transparency, reference transparency and background transparency corresponding to the current instance.
16. The image processing method according to claim 15, characterized in that: The optimizing the image features corresponding to the current instance based on the image features corresponding to the remaining instances to obtain the optimized image features corresponding to the current instance includes: Decomposing the image features corresponding to each instance of the at least one instance into target sub-features, reference sub-features and background sub-features; Based on the target sub-features, reference sub-features and background sub-features corresponding to the remaining instances, feature fusion is performed on the target sub-features, reference sub-features and background sub-features corresponding to the current instance to obtain optimized target sub-features, reference sub-features and background sub-features corresponding to the current instance; The optimized target sub-feature, reference sub-feature and background sub-feature corresponding to the current instance are aggregated to obtain the optimized image feature corresponding to the current instance.
17. The image processing method according to claim 16, characterized in that: The step of decomposing the image features corresponding to each instance of the at least one instance into target sub-features, reference sub-features and background sub-features comprises: Combining the image feature corresponding to each instance with the target transparency of each instance to obtain the target sub-feature of each instance; Combining the image feature corresponding to each instance and the reference transparency of each instance to obtain the reference sub-feature of each instance; The image feature corresponding to each instance and the background transparency of each instance are combined to obtain the background sub-feature of each instance.
18. The image processing method according to claim 16, characterized in that: The performing feature fusion on the target sub-feature, the reference sub-feature and the background sub-feature corresponding to the current instance based on the target sub-feature, the reference sub-feature and the background sub-feature corresponding to the remaining instances includes: Based on the target sub-features and reference sub-features corresponding to each instance, feature reduction is performed on the target sub-feature corresponding to the current instance to obtain an optimized target sub-feature corresponding to the current instance; Based on the reference sub-feature corresponding to the current instance and the target sub-features corresponding to the remaining instances, feature reduction is performed on the reference sub-feature corresponding to the current instance to obtain an optimized reference sub-feature corresponding to the current instance; Based on the background sub-features corresponding to each instance, feature reduction is performed on the background sub-features corresponding to the current instance to obtain optimized background sub-features corresponding to the current instance.
19. A training device for an image processing model, characterized in that: include: A first image acquisition unit is configured to: acquire a training image, wherein the training image includes at least one instance, and each instance is marked with a true transparency; A first mask acquisition unit is configured to: acquire a mask of each instance of the at least one instance; The first model prediction unit is configured to: for each instance, based on the mask of the at least one instance, generate a target mask, a reference mask and a background mask of the current instance, and input the training image, the target mask, the reference mask and the background mask into the image processing model to obtain the target transparency, the reference transparency and the background transparency of the current instance, wherein the target mask represents the mask of the current instance, the reference mask represents the masks of the remaining instances except the current instance, the background mask represents the mask of the remaining parts of the training image except the at least one instance, the target transparency is the transparency corresponding to the target mask, the reference transparency is the transparency corresponding to the reference mask, and the background transparency is the transparency corresponding to the background mask; A loss calculation unit, configured to: calculate a loss based on the target transparency, the reference transparency, the background transparency and the real transparency of each instance; A parameter adjustment unit is configured to train the image processing model by adjusting the parameters of the image processing model based on the loss.
20. The training device according to claim 19, characterized in that The image processing model includes a feature extraction module and a transparency prediction module; The first model prediction unit is configured as follows: Inputting the training image, the target mask, the reference mask and the background mask into the feature extraction module to obtain image features corresponding to the current instance; The image features are input into the transparency prediction module to obtain the target transparency, reference transparency and background transparency of the current instance.
21. The training device according to claim 19 or 20, characterized in that The first optimization unit is configured to: For each instance, the target transparency, reference transparency and background transparency of the remaining instances are obtained, and the target transparency, reference transparency and background transparency of the current instance are optimized based on the target transparency, reference transparency and background transparency of the remaining instances to obtain the optimized target transparency, reference transparency and background transparency of the current instance.
22. The training device according to claim 21, characterized in that The loss calculation unit is configured as: Calculating a first loss based on the target transparency, the reference transparency, the background transparency, and the true transparency of each instance; Calculating a second loss based on the optimized target transparency, the reference transparency, the background transparency, and the true transparency of each instance; The loss is determined based on the first loss and the second loss.
23. The training device according to claim 21, characterized in that The first optimization unit is configured as: For each instance, do the following: Acquire an image feature corresponding to the at least one instance; Optimizing the image features corresponding to the current instance based on the image features corresponding to the remaining instances to obtain optimized image features corresponding to the current instance; The optimized image features corresponding to the current instance are input into the transparency prediction module to obtain the optimized target transparency, reference transparency and background transparency corresponding to the current instance.
24. The training device according to claim 23, characterized in that The first optimization unit is configured as: Decomposing the image features corresponding to each instance of the at least one instance into target sub-features, reference sub-features and background sub-features; Based on the target sub-features, reference sub-features and background sub-features corresponding to the remaining instances, feature fusion is performed on the target sub-features, reference sub-features and background sub-features corresponding to the current instance to obtain optimized target sub-features, reference sub-features and background sub-features corresponding to the current instance; The optimized target sub-feature, reference sub-feature and background sub-feature corresponding to the current instance are aggregated to obtain the optimized image feature corresponding to the current instance.
25. The training device according to claim 24, characterized in that The first optimization unit is configured as: Combining the image feature corresponding to each instance with the target transparency of each instance to obtain the target sub-feature of each instance; Combining the image feature corresponding to each instance and the reference transparency of each instance to obtain the reference sub-feature of each instance; The image feature corresponding to each instance and the background transparency of each instance are combined to obtain the background sub-feature of each instance.
26. The training device according to claim 24, characterized in that The first optimization unit is configured as: Based on the target sub-features and reference sub-features corresponding to each instance, feature reduction is performed on the target sub-feature corresponding to the current instance to obtain an optimized target sub-feature corresponding to the current instance; Based on the reference sub-feature corresponding to the current instance and the target sub-features corresponding to the remaining instances, feature reduction is performed on the reference sub-feature corresponding to the current instance to obtain an optimized reference sub-feature corresponding to the current instance; Based on the background sub-features corresponding to each instance, feature reduction is performed on the background sub-features corresponding to the current instance to obtain optimized background sub-features corresponding to the current instance.
27. The training device according to claim 19, characterized in that The loss calculation unit is configured as: determining a first minimum absolute deviation among the target transparency, the reference transparency, the background transparency, and the true transparency of each instance; Based on the target transparency, the reference transparency, the background transparency and the true transparency of each instance, smoothness constraints are imposed on the parameters of the image processing model to obtain the first Laplace loss of each instance; Based on the target synthesis part, the reference synthesis part, the background synthesis part and the training image of each instance, a first synthesis constraint loss of each instance is obtained, wherein the target synthesis part is obtained based on the target transparency and the target instance foreground image, the reference synthesis part is obtained based on the reference transparency and the reference instance foreground image, the background synthesis part is obtained based on the background transparency and the background image, and the target instance foreground image, the reference instance foreground image and the background image are all determined based on the real transparency and the training image; Based on the target transparency, the reference transparency and the background transparency of each instance, obtaining a first transparency constraint loss of each instance; The loss is determined based on the first minimum absolute deviation of each instance, the first Laplace loss, the first synthesis constraint loss, and the first transparency constraint loss.
28. The training device according to claim 22, characterized in that The loss calculation unit is configured as: Determining a first minimum absolute deviation among the target transparency, the reference transparency, the background transparency, and the true transparency of each instance, and determining a second minimum absolute deviation among the optimized target transparency, the reference transparency, the background transparency, and the true transparency of each instance; Based on the target transparency, the reference transparency, the background transparency and the true transparency of each instance, smooth constraints are imposed on the parameters of the image processing model to obtain a first Laplace loss of each instance, and based on the optimized target transparency, the reference transparency, the background transparency and the true transparency of each instance, smooth constraints are imposed on the parameters of the image processing model to obtain a second Laplace loss of each instance; Based on the target synthesis part, the reference synthesis part, the background synthesis part and the training image of each instance, a first synthesis constraint loss of each instance is obtained, and based on the optimized target synthesis part, the optimized reference synthesis part, the optimized background synthesis part and the training image of each instance, a second synthesis constraint loss of each instance is obtained, wherein the target synthesis part is obtained based on the target transparency and the target instance foreground image, the reference synthesis part is obtained based on the reference transparency and the reference instance foreground image, the background synthesis part is obtained based on the background transparency and the background image, the target instance foreground image, the reference instance foreground image and the background image are all determined based on the real transparency and the training image, the optimized target synthesis part is obtained based on the optimized target transparency and the target instance foreground image, the optimized reference synthesis part is obtained based on the optimized reference transparency and the reference instance foreground image, and the optimized background synthesis part is obtained based on the optimized background transparency and the background image; Based on the target transparency, the reference transparency and the background transparency of each instance, a first transparency constraint loss of each instance is obtained, and based on the optimized target transparency, the reference transparency and the background transparency of each instance, a second transparency constraint loss of each instance is obtained; determining the first loss based on the first minimum absolute deviation, the first Laplace loss, the first composite constraint loss, and the first transparency constraint loss of each of the instances, and determining the second loss based on the second minimum absolute deviation, the second Laplace loss, the second composite constraint loss, and the second transparency constraint loss of each of the instances; The loss is determined based on the first loss and the second loss.
29. An image processing device, characterized in that: include: A second image acquisition unit is configured to: acquire an image to be processed, wherein the image to be processed includes at least one instance; A second mask acquisition unit is configured to: acquire a mask of each instance of the at least one instance; The second model prediction unit is configured to: for each instance, based on the mask of the at least one instance, generate a target mask, a reference mask and a background mask of the current instance, and input the image to be processed, the target mask, the reference mask and the background mask into the image processing model to obtain the target transparency, the reference transparency and the background transparency of the current instance, wherein the target mask represents the mask of the current instance, the reference mask represents the masks of the remaining instances except the current instance, the background mask represents the mask of the remaining parts of the image to be processed except the at least one instance, the target transparency is the transparency corresponding to the target mask, the reference transparency is the transparency corresponding to the reference mask, and the background transparency is the transparency corresponding to the background mask; The instance extraction unit is configured to extract each instance from the to-be-processed image based on the target transparency, the reference transparency and the background transparency of each instance.
30. The image processing device according to claim 29, wherein: The image processing model includes a feature extraction module and a transparency prediction module; The second model prediction unit is configured as: Inputting the image to be processed, the target mask, the reference mask and the background mask into the feature extraction module to obtain image features corresponding to the current instance; The image features are input into the transparency prediction module to obtain the target transparency, reference transparency and background transparency of the current instance.
31. The image processing device according to claim 29 or 30, characterized in that: Also included is a second optimization unit configured to: For each instance, the target transparency, reference transparency and background transparency of the remaining instances are obtained, and the target transparency, reference transparency and background transparency of the current instance are optimized based on the target transparency, reference transparency and background transparency of the remaining instances to obtain the optimized target transparency, reference transparency and background transparency of the current instance.
32. The image processing device according to claim 31, characterized in that: The instance extraction unit is configured as follows: Each instance is extracted from the image to be processed based on the optimized target transparency, reference transparency and background transparency of each instance.
33. The image processing device according to claim 31, characterized in that: The second optimization unit is configured as follows: For each instance, do the following: Acquire an image feature corresponding to the at least one instance; Optimizing the image features corresponding to the current instance based on the image features corresponding to the remaining instances to obtain optimized image features corresponding to the current instance; The optimized image features corresponding to the current instance are input into the transparency prediction module to obtain the optimized target transparency, reference transparency and background transparency corresponding to the current instance.
34. The image processing device according to claim 33, characterized in that: The second optimization unit is configured as follows: Decomposing the image features corresponding to each instance of the at least one instance into target sub-features, reference sub-features and background sub-features; Based on the target sub-features, reference sub-features and background sub-features corresponding to the remaining instances, feature fusion is performed on the target sub-features, reference sub-features and background sub-features corresponding to the current instance to obtain optimized target sub-features, reference sub-features and background sub-features corresponding to the current instance; The optimized target sub-feature, reference sub-feature and background sub-feature corresponding to the current instance are aggregated to obtain the optimized image feature corresponding to the current instance.
35. The image processing device according to claim 34, characterized in that: The second optimization unit is configured as follows: Combining the image feature corresponding to each instance with the target transparency of each instance to obtain the target sub-feature of each instance; Combining the image feature corresponding to each instance and the reference transparency of each instance to obtain the reference sub-feature of each instance; The image feature corresponding to each instance and the background transparency of each instance are combined to obtain the background sub-feature of each instance.
36. The image processing device according to claim 34, characterized in that: The second optimization unit is configured as follows: Based on the target sub-features and reference sub-features corresponding to each instance, feature reduction is performed on the target sub-feature corresponding to the current instance to obtain an optimized target sub-feature corresponding to the current instance; Based on the reference sub-feature corresponding to the current instance and the target sub-features corresponding to the remaining instances, feature reduction is performed on the reference sub-feature corresponding to the current instance to obtain an optimized reference sub-feature corresponding to the current instance; Based on the background sub-features corresponding to each instance, feature reduction is performed on the background sub-features corresponding to the current instance to obtain optimized background sub-features corresponding to the current instance.
37. An electronic device, characterized in that: include: at least one processor; at least one memory storing computer executable instructions, Wherein, when the computer executable instructions are executed by the at least one processor, they prompt the at least one processor to execute the training method of the image processing model as described in any one of claims 1 to 10 or the image processing method as described in any one of claims 11 to 18.
38. A computer-readable storage medium storing instructions, characterized in that: When the instructions are executed by at least one processor, the at least one processor is prompted to execute the training method of the image processing model as described in any one of claims 1 to 10 or the image processing method as described in any one of claims 11 to 18.
39. A computer program product comprising computer instructions, characterized in that: When the computer instructions are executed by at least one processor, the training method of the image processing model as described in any one of claims 1 to 10 or the image processing method as described in any one of claims 11 to 18 is implemented.