Model generation method, image processing method, device, and electronic device
By training the matting model in multiple stages and combining encoding, decoding and classification networks, the problem of insufficient resolution adaptability of existing matting models is solved, and the flexibility and diversity of output results are achieved.
Patent Information
- Application Number
- CN202310552962.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-05-16
- Publication Date
- 2025-11-11
- Estimated Expiration
- 2043-05-16
AI Technical Summary
Existing image matting models lack flexibility and are unable to meet the resolution requirements of different scenarios and tasks.
The matting model is trained in multiple stages using a training dataset and loss function. First, a low-resolution matting model is generated. Then, an encoding network, a decoding network, and a classification network are combined. Finally, the resolution of the output result is determined by a judgment module, so that the output results corresponding to different classification results have different resolutions.
It improves the flexibility of the image matting model, enabling it to output results at different resolutions for different types of images, meeting the needs of various scenarios and tasks.
Smart Images

Figure CN116758108B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer technology, and more specifically, to a model generation method, an image processing method, an apparatus, and an electronic device. Background Technology
[0002] With the continuous development of image processing, image matting technology has begun to be widely used. One approach involves training a matting model using a matting dataset to obtain the model, which can then be used to perform image matting. However, the flexibility of this matting model still needs improvement. Summary of the Invention
[0003] In view of the above problems, this application proposes a model generation method, an image processing method, an apparatus, and an electronic device to improve the above problems.
[0004] In a first aspect, this application provides a model generation method, the method comprising: training a matting model to be trained based on a first training dataset and a first loss function to obtain a low-resolution matting model, wherein the first training dataset includes multiple subject images and first annotation information for each of the multiple subject images; training a classification model to be trained based on a second training dataset and a second loss function to obtain a target classification model, wherein the second training dataset includes multiple subject images identical to the first training dataset and second annotation information for each of the multiple subject images, the classification model to be trained includes a classification network and an encoding network of the low-resolution matting model; and training the classification matting model to be trained based on the first training dataset and a third loss function to obtain a target classification matting model, wherein the classification matting model to be trained includes the encoding network, the decoding network of the low-resolution matting model, the classification network, and a judgment module, the judgment module being used to determine the output result of the target classification matting model based on the classification result of the classification network, wherein the resolution of the output result corresponding to different classification results is different.
[0005] Secondly, this application provides an image processing method, the method comprising: acquiring an image to be cut out; inputting the image to be cut out into a target classification cutout model to obtain a cutout result of the image to be cut out, wherein the target classification cutout model is obtained based on the above method.
[0006] Thirdly, this application provides a model generation apparatus, comprising: a model generation unit, configured to train a matting model to be trained based on a first training dataset and a first loss function to obtain a low-resolution matting model, wherein the first training dataset includes multiple subject images and first annotation information for each of the multiple subject images; to train a classification model to be trained based on a second training dataset and a second loss function to obtain a target classification model, wherein the second training dataset includes multiple subject images identical to the first training dataset and second annotation information for each of the multiple subject images, and the classification model to be trained includes a classification network and an encoding network of the low-resolution matting model; and to train a classification matting model to be trained based on the first training dataset and a third loss function to obtain a target classification matting model, wherein the classification matting model to be trained includes the encoding network, the decoding network of the low-resolution matting model, the classification network, and a judgment module, wherein the judgment module is configured to determine the output result of the target classification matting model based on the classification result of the classification network, wherein the resolution of the output result corresponding to different classification results is different.
[0007] Fourthly, this application provides an image processing apparatus, the apparatus comprising: an image to be cut out acquisition unit, configured to acquire an image to be cut out; and a cutout result acquisition unit, configured to input the image to be cut out into a target classification cutout model to obtain a cutout result of the image to be cut out, wherein the target classification cutout model is obtained based on the above method.
[0008] Fifthly, this application provides an electronic device including one or more processors and a memory; one or more programs are stored in the memory and configured to be executed by the one or more processors, the one or more programs being configured to perform the methods described above.
[0009] Sixthly, this application provides a computer-readable storage medium storing program code, wherein the above-described method is executed when the program code is run.
[0010] This application provides a model generation method, image processing method, apparatus, electronic device, and storage medium. After training a matting model to be trained based on multiple subject images, a first training dataset containing first annotation information for each of the multiple subject images, and a first loss function, to obtain a low-resolution matting model, a classification model to be trained is then trained based on multiple subject images identical to the first training dataset, a second training dataset containing second annotation information for each of the multiple subject images, and a second loss function. This yields a target classification model. Finally, a classification matting model to be trained, comprising the encoding network, the decoding network of the low-resolution matting model, the classification network, and a judgment module, is trained based on the first training dataset and a third loss function to obtain a target classification matting model. The judgment module is used to determine the output result of the target classification matting model based on the classification result of the classification network, and different classification results correspond to output results with different resolutions. The above method allows for the following steps: First, a low-resolution matting model is trained using a training dataset and a first loss function. Then, a classification model containing the low-resolution matting model is trained using the training dataset and a second loss function to obtain a target classification model. Finally, a classification matting model containing the low-resolution matting model, its encoding and decoding networks, and the target classification model's classification network is trained using the training dataset and a third loss function to obtain a target classification matting model. This allows the target classification matting model to output different resolutions for different types of images, thereby improving its matting flexibility and ensuring that its output meets the resolution requirements of different scenarios and tasks. Attached Figure Description
[0011] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0012] Figure 1 A flowchart of a model generation method proposed in an embodiment of this application is shown;
[0013] Figure 2 A schematic diagram of a matting model to be trained according to this application is shown;
[0014] Figure 3 A schematic diagram of a classification model to be trained proposed in this application is shown;
[0015] Figure 4 This application shows Figure 1 A flowchart of one embodiment of S130;
[0016] Figure 5 A schematic diagram of a classification matting model to be trained proposed in this application is shown;
[0017] Figure 6 A flowchart of an image processing method proposed in an embodiment of this application is shown;
[0018] Figure 7 A structural block diagram of a model generation apparatus according to an embodiment of this application is shown;
[0019] Figure 8 A structural block diagram of an image processing apparatus according to an embodiment of this application is shown;
[0020] Figure 9 A structural block diagram of an electronic device proposed in this application is shown;
[0021] Figure 10 It is a storage unit in this application embodiment for storing or carrying program code that implements the model generation method and image processing method according to the embodiments of this application. Detailed Implementation
[0022] The technical solutions in the embodiments of this application will be clearly and completely described below with reference to the appendices. All other embodiments obtained by those skilled in the art based on the embodiments of this application without inventive effort are within the scope of protection of this application.
[0023] With the continuous development of artificial intelligence technology, image matting has become an important task in the field of computer vision. Image matting technology can be used to divide an image into foreground and background parts, and extract the content of the foreground part. The content of the foreground part can be called the subject, which is the content to be extracted. The subject can vary depending on the specific matting task. For example, an image may contain people, animals, trees, etc. When the matting task is portrait matting, the subject can be a person; when the matting task is animal matting, the subject can be an animal.
[0024] However, the inventors found in their research that the flexibility of the image cutout method still needs to be improved.
[0025] Therefore, the inventors have proposed a model generation method, image processing method, apparatus, and electronic device in this application. After training a matting model to be trained based on multiple subject images, a first training dataset containing first annotation information for each of the multiple subject images, and a first loss function, to obtain a low-resolution matting model, a classification model to be trained is then trained based on multiple subject images identical to the first training dataset, a second training dataset containing second annotation information for each of the multiple subject images, and a second loss function, to obtain a target classification model. Finally, a classification matting model to be trained, comprising the encoding network and the encoding network of the low-resolution matting model, is then trained based on the first training dataset and a third loss function, to obtain a target classification matting model. The judgment module is used to determine the output result of the target classification matting model based on the classification result of the classification network, and different classification results correspond to output results with different resolutions. The above method allows for the following steps: First, a low-resolution matting model is trained using a training dataset and a first loss function. Then, a classification model containing the low-resolution matting model is trained using the training dataset and a second loss function to obtain a target classification model. Finally, a classification matting model containing the low-resolution matting model, its encoding and decoding networks, and the target classification model's classification network is trained using the training dataset and a third loss function to obtain a target classification matting model. This allows the target classification matting model to output different resolutions for different types of images, thereby improving its matting flexibility and ensuring that its output meets the resolution requirements of different scenarios and tasks.
[0026] The embodiments of this application will now be described in conjunction with the accompanying drawings.
[0027] Please see Figure 1 This application provides a model generation method, the method comprising:
[0028] S110: Train the matting model to be trained based on the first training dataset and the first loss function to obtain a low-resolution matting model. The first training dataset includes multiple subject images and the first annotation information of each of the multiple subject images.
[0029] The subject image can refer to an image containing a subject. The first annotation information can include multiple ground truth mask images, which can be grayscale images. Grayscale images can refer to images with pixel values ranging from 0 to 255, so that the grayscale images can represent the pixel value of each pixel of the subject in the corresponding subject image, and distinguish the subject and the background in the corresponding subject image.
[0030] In this embodiment, the matting model to be trained can be a deep learning model for matting. The matting model to be trained can include an encoding network and a decoding network. The encoding network can include multiple encoders, and the decoding network can include multiple decoders. For example, the matting model to be trained can be a U2NET model, etc.
[0031] As a way, such as Figure 2 As shown, multiple ground truth mask images can be down-resolution processed to obtain multiple low-resolution ground truth mask images; multiple subject images are input into the matting model to be trained to obtain multiple first prediction mask images that correspond one-to-one with the multiple subject images; the matting model to be trained is trained based on the multiple low-resolution ground truth mask images, the multiple first prediction mask images, and the first loss function to obtain a low-resolution matting model. The first loss function can be used to reduce the pixel value difference between each pixel in the multiple low-resolution ground truth mask images and the multiple first prediction mask images.
[0032] The first loss function can be a regression loss function, such as L1 loss, and the formula for calculating the first loss function is as follows:
[0033] L1=1 / n∑|x i -y i |
[0034] Where n can represent the number of main images, x i The first predicted mask image can be represented as y. i It can represent low-resolution ground truth mask images.
[0035] Optionally, since the resolution of the low-resolution ground truth mask image obtained after inputting the subject image into the matting model to be trained will be smaller than that of the ground truth mask image, in order to facilitate the calculation of the loss function, multiple ground truth mask images can be downsampled (for example, using cv2.pyrDown() in OpenCV) to obtain multiple low-resolution ground truth mask images, so that the resolution of the low-resolution ground truth mask images can be the same as the resolution of the first prediction mask image.
[0036] Optionally, the first training dataset can be obtained based on an open-source matting dataset downloaded from the internet and / or a matting dataset constructed according to task requirements. For example, the first training dataset can have approximately 500,000 subject images, which can come from an open-source matting dataset or from a matting dataset constructed according to task requirements.
[0037] S120: Train the classification model to be trained based on the second training dataset and the second loss function to obtain the target classification model. The second training dataset includes multiple subject images that are the same as the first training dataset, as well as the second annotation information of each of the multiple subject images. The classification model to be trained includes a classification network and the encoding network of the low-resolution matting model.
[0038] The second annotation information may include the real category labels corresponding to each of the multiple subject images, such as people, animals, plants, etc.
[0039] As a way, such as Figure 3 As shown, multiple subject images can be input into the encoding network of a low-resolution matting model to obtain multiple feature maps that correspond one-to-one with the multiple subject images; the multiple feature maps are then input into a classification network to obtain prediction information corresponding to each of the multiple subject images; based on the true class labels corresponding to each of the multiple subject images, the prediction information corresponding to each of the multiple subject images, and a second loss function, the classification network is trained to obtain a target classification model. The second loss function can be used to reduce the difference between the true class labels corresponding to each of the multiple subject images and the prediction information corresponding to each of the multiple subject images.
[0040] The classification network may include at least one convolutional layer and at least one fully connected layer. For example, a classification network may include five 3x3 convolutional layers and one fully connected (FC) layer.
[0041] The prediction information may include a prediction probability array and a prediction category label. The prediction probability array for each subject image may contain the probability value that the subject image is predicted to be the corresponding category.
[0042] Optionally, multiple feature maps are input into the classification network to obtain prediction probability arrays corresponding to each of the multiple subject images. The prediction probability array of each subject image may contain the probability value of the subject image being predicted as the corresponding category, and the category with the highest probability value is taken as the predicted category of the subject image.
[0043] For example, the true category label of each subject image can be 0 or 1, where 1 can represent the category of human and 0 can represent the category of non-human. The predicted probability array of the subject image can be [0.85, 0.15], which means that the probability of the subject image being predicted as human is 0.85 and the probability of it being predicted as non-human is 0.15. Furthermore, it can mean that the predicted category label of the subject image can be 1.
[0044] The second loss function can be a classification loss function, such as the cross-entropy loss function. When the true class labels of multiple subject images are in total two classes, the formula for calculating the second loss function can be:
[0045]
[0046] Where n can represent the number of main images, y i p can represent the true category label of the i-th subject image. i It can represent the probability of predicting the true category label.
[0047] When multiple subject images have at least three true class labels, the true class label of each subject image can be one-hot encoded. For example, the true class of each subject image can be human, cat, or dog. When the true class label of a subject image is [0 0 1], it can represent that the true class of the subject image is dog. In this case, the formula for calculating the second loss function can be:
[0048]
[0049] Where n can represent the number of subject images, k can represent the total number of true category labels for the subject images, and y i,j p can represent the encoded value of the i-th subject image under label j. i,j It can represent the probability that the predicted category label of the i-th subject image is j.
[0050] It's important to note that during the training of the classification model, the network parameters (such as weights) of the encoding network remain unchanged; only the parameters of the classification network are updated. In other words, when training the classification model, the encoding network needs to be frozen, and only the classification network is trained.
[0051] In this embodiment of the application, by freezing the encoding network and training only the classification network, the encoding network in the target classification model can still retain the ability to extract features from the subject image during the low-resolution matting model period, while enabling the classification network to accurately predict the category of the subject image.
[0052] S130: The training model for classification matting is trained based on the first training dataset and the third loss function to obtain the target classification matting model. The training model for classification matting includes the encoding network, the decoding network of the low-resolution matting model, the classification network, and the judgment module. The judgment module is used to determine the output result of the target classification matting model based on the classification result of the classification network. The resolution of the output result corresponding to different classification results is different.
[0053] The classification matting model to be trained may further include an upsampling network, which can be a network used to improve image resolution. In the embodiments of this application, the upsampling network can have various implementations.
[0054] One approach is to use a PixelShuffle (Sub-Pixel Convolutional Neural Network) network for upsampling. PixelShuffle divides a low-resolution pixel into r×r parts, by default, based on the r values of the corresponding pixel locations in the feature map. 2 Each feature pixel is composed of several feature pixels, forming a low-resolution pixel. During this composition process, the best upsampling effect can be achieved by continuously optimizing the weights of each combination. The PixelShuffle network can first perform multiple convolution operations on a low-resolution image (Input) of size H×W×C to obtain an image of size H×W×(C×r). 2 The feature map is obtained by performing a shuffle transformation on the feature map to obtain a super-resolution image (output) with size H×(W×r)×(C×r).
[0055] Alternatively, an upsampling network can include at least one deconvolutional (transposed convolutional) layer.
[0056] As another approach, the upsampling network may include at least one unpooling layer.
[0057] As a way, such as Figure 4 As shown, the target classification matting model is trained based on the first training dataset and the third loss function to obtain the target classification matting model, including:
[0058] S131: Input the multiple subject images into the encoding network to obtain multiple feature maps that correspond one-to-one with the multiple subject images.
[0059] One approach is to input multiple subject images into an encoding network to obtain multiple feature maps that correspond one-to-one with the multiple subject images.
[0060] It should be noted that the network parameters (such as weights) of the encoding network do not change during the training process of the classification matting model. In other words, the encoding network needs to be frozen during the training of the classification matting model.
[0061] In this embodiment of the application, by freezing the encoding network and training only the classification network, the encoding network in the target classification model can still retain the ability to extract features from the subject image during the low-resolution matting model period, while enabling the classification network to accurately predict the category of the subject image.
[0062] S132: Input the multiple feature maps into the classification network to obtain the predicted categories corresponding to each of the multiple subject images.
[0063] As a way, such as Figure 5 As shown, multiple feature maps can be input into the classification network to obtain prediction probability arrays corresponding to each subject image. The prediction probability array of each subject image can contain the probability value of the subject image being predicted as the corresponding category. The category with the highest probability value is taken as the predicted category of the subject image.
[0064] It should be noted that the network parameters (such as weights) of the classification network will not change during the training process of the classification matting model. In other words, the classification network needs to be frozen during the training of the classification matting model.
[0065] S133: Input the multiple feature maps into the decoding network to obtain multiple second prediction mask images that correspond one-to-one with the multiple main images.
[0066] One approach is to input multiple feature maps into the decoding network to obtain multiple second prediction mask images that correspond one-to-one with multiple main images.
[0067] It should be noted that the network parameters (such as weights) of the decoding network remain unchanged during the training of the classification matting model. In other words, the decoding network needs to be frozen during the training of the classification matting model.
[0068] Furthermore, it should be noted that there is no explicit order between steps S132 and S133; they can be executed simultaneously or at different times.
[0069] S134: Input the predicted categories corresponding to each of the multiple subject images and the multiple second prediction mask images into the judgment module to obtain the target subject image, wherein the target subject image is the subject image whose corresponding predicted category is determined by the judgment module to belong to the target category.
[0070] Here, the target category can refer to the category of the corresponding second predictive mask image that needs to be processed to increase resolution. For example, the target category can be set to "person". Increasing the resolution of the second predictive mask image of the target category can improve the image matting effect of the subject image of the target category. For example, the edges of the subject will be clearer, and more details will be extracted.
[0071] The judgment module can be implemented through conditional statements. When the predicted category of the subject image is the target category, the subject image can be determined to be the target subject image and input into the upsampling network. When the predicted category of the subject image is not the target category, the subject image can be determined not to be the target subject image and output directly.
[0072] One approach is to input the predicted categories corresponding to multiple subject images and multiple second prediction mask images into the judgment module, and use the subject image whose predicted category is the same as the target category as the target subject image.
[0073] For example, if the target category can be "human" and the predicted category of the subject image can be "human", then the subject image is the target subject image.
[0074] In the embodiments of this application, the target category may include multiple target subcategories, the upsampling network may include multiple upsampling subnetworks, each upsampling subnetwork may be used to output the output result of the corresponding target subcategory, and the resolution of the output result of each target subcategory is different.
[0075] In one approach, the predicted categories corresponding to multiple subject images and multiple second prediction mask images can be input into the judgment module to obtain the target subject image. The target subject image can be a subject image whose corresponding predicted category is determined by the judgment module to belong to any one of the multiple target subcategories.
[0076] S135: Input the second prediction mask image corresponding to the target subject image into the upsampling network to obtain a third prediction mask image, wherein the resolution of the third prediction mask image is higher than that of the second prediction mask image.
[0077] As one approach, when the predicted categories of the target subject images are the same, that is, when the target category includes only one category, the second predicted mask image corresponding to the target subject image can be input into the upsampling network to obtain the third predicted mask image.
[0078] As one approach, when the predicted categories of the target subject image are different, that is, when the target category includes multiple categories, the second prediction mask image corresponding to the target subject image can be input into the corresponding upsampling sub-network to obtain the third prediction mask image.
[0079] For example, multiple target subcategories may include people and animals. The upsampling subnetwork corresponding to people can increase the resolution of the input image by 4 times, and the upsampling subnetwork corresponding to animals can increase the resolution of the input image by 2 times. Thus, the third prediction mask image of the target subject image with the predicted category of people has a resolution of 4 times that of the corresponding second prediction mask image, and the third prediction mask image of the target subject image with the predicted category of animals has a resolution of 2 times that of the corresponding second prediction mask image.
[0080] As one approach, when the subject image is not the target image, the second predicted mask image can be directly used as the third mask image.
[0081] S136: The upsampling network is trained based on the third predicted mask image, the ground truth mask image corresponding to the third predicted mask image, and the third loss function to obtain the target classification matting model, wherein the third loss function is used to reduce the pixel value difference between each pixel in the third predicted mask image and the corresponding ground truth mask image.
[0082] One approach is to first obtain a target mask image, which can be an image with the same resolution as the third predicted mask image, obtained based on the ground truth mask image corresponding to the third predicted mask image. The upsampling network is then trained based on the third predicted mask image, the target mask image, and the third loss function to obtain a target classification matting model.
[0083] The third loss function can be a regression loss function, such as MSE (Mean Square Error), and the formula for calculating the first loss function can be:
[0084] K3=1 / n∑(x i -y i ) 2
[0085] Where n can represent the number of target subject images, x i It can represent the third prediction mask image, y i It can represent a target mask image.
[0086] Optionally, the ground truth mask image corresponding to the third predicted mask image can be down-resolution or up-resolution to obtain the target mask image.
[0087] In this embodiment, by freezing the encoding network, decoding network, and classification network, and training only the upsampling network, the encoding network in the target classification matting model can still retain the ability to extract features from the subject image during the low-resolution matting model period; the decoding network can still retain the ability to extract features from the subject image during the low-resolution matting model period; the classification network can still retain the ability to classify the subject image during the target classification model period; and the upsampling network can improve the resolution of the third prediction mask image of the target category.
[0088] Furthermore, by setting up an upsampling network, different types of subject images can obtain third prediction mask images of different resolutions, thereby meeting the image matting requirements of different scenes and tasks.
[0089] This embodiment provides a model generation method. First, a low-resolution matting model is trained using a first training dataset and a first loss function, comprising multiple subject images, each with its own first annotation information, and a first training dataset. Then, a target classification model is trained using a second training dataset and a second loss function, comprising a classification network and an encoding network of the low-resolution matting model, based on the same multiple subject images as the first training dataset and their respective second annotation information. Finally, a target classification matting model is trained using the first training dataset and a third loss function, comprising the encoding network, a decoding network of the low-resolution matting model, the classification network, and a judgment module. The judgment module determines the output of the target classification matting model based on the classification results of the classification network, and different classification results correspond to different resolutions in the output. The above method allows for the following steps: First, a low-resolution matting model is trained using a training dataset and a first loss function. Then, a classification model containing the low-resolution matting model is trained using the training dataset and a second loss function to obtain a target classification model. Finally, a classification matting model containing the low-resolution matting model, its encoding and decoding networks, and the target classification model's classification network is trained using the training dataset and a third loss function to obtain a target classification matting model. This allows the target classification matting model to output different resolutions for different types of images, thereby improving its matting flexibility and ensuring that its output meets the resolution requirements of different scenarios and tasks.
[0090] Please see Figure 6 This application provides an image processing method, the method comprising:
[0091] S210: Obtain the image to be cut out.
[0092] One method is to acquire the image to be cut out using an image acquisition device (such as a camera, webcam, or mobile phone).
[0093] S220: Input the image to be matted into the target classification matting model to obtain the matting result of the image to be matted, wherein the target classification matting model is obtained based on any one of the methods described in claims 1-6.
[0094] One approach is to input the image to be matted into a target classification matting model to obtain a mask image of the image to be matted; based on the image to be matted and the mask image, the matting result is obtained.
[0095] The target classification matting model can include an encoding network, a decoding network, a classification network, a judgment module, and an upsampling network.
[0096] Optionally, multiple subject images can be input into an encoding network to obtain multiple feature maps corresponding one-to-one with the multiple subject images; the multiple feature maps can be input into a classification network to obtain the predicted category corresponding to each of the multiple subject images; the multiple feature maps can be input into a decoding network to obtain multiple reference prediction mask images corresponding one-to-one with the multiple subject images; the prediction information corresponding to each of the multiple subject images and the multiple reference prediction mask images can be input into a judgment module. If the judgment module determines that the predicted category is the target category, the reference prediction mask image can be input into an upsampling network to obtain a mask image; if the judgment module determines that the predicted category is not the target category, the reference prediction mask image can be used as the mask image.
[0097] Optionally, the pixel values at the same location in the mask image and the image to be cut out can be multiplied to obtain the cutout result. For example, the cutout result can be as follows: Figure 3 As shown.
[0098] This embodiment provides an image processing method that, through the above-described manner, allows the image to be matted to be input into a target matting model that includes a judgment module and an upsampling network. This enables the output of different resolutions for different types of images to be matted, improving the matting flexibility of the target classification matting model so that the output of the target classification matting model can meet the resolution requirements of different scenarios and tasks.
[0099] Please see Figure 7 This application provides a model generation apparatus 600, the apparatus 600 comprising:
[0100] The model generation unit 610 is used to train a matting model to be trained based on a first training dataset and a first loss function to obtain a low-resolution matting model. The first training dataset includes multiple subject images and first annotation information for each of the multiple subject images. It is also used to train a classification model to be trained based on a second training dataset and a second loss function to obtain a target classification model. The second training dataset includes multiple subject images identical to those in the first training dataset and second annotation information for each of the multiple subject images. The classification model to be trained includes a classification network and an encoding network of the low-resolution matting model. Finally, it is used to train a classification matting model to be trained based on the first training dataset and a third loss function to obtain a target classification matting model. The classification matting model to be trained includes the encoding network, the decoding network of the low-resolution matting model, the classification network, and a judgment module. The judgment module is used to determine the output result of the target classification matting model based on the classification result of the classification network. Different classification results correspond to output results with different resolutions.
[0101] In one approach, the first annotation information includes multiple ground truth mask images. The model generation unit 610 is specifically used to down-resolution the multiple ground truth mask images to obtain multiple low-resolution ground truth mask images; input the multiple subject images into the matting model to be trained to obtain multiple first prediction mask images that correspond one-to-one with the multiple subject images; train the matting model to be trained based on the multiple low-resolution ground truth mask images, the multiple first prediction mask images, and the first loss function to obtain the low-resolution matting model, wherein the first loss function is used to reduce the pixel value difference between each pixel of the multiple low-resolution ground truth mask images and the multiple first prediction mask images.
[0102] In one approach, the second annotation information includes the ground truth class labels corresponding to each of the multiple subject images. The model generation unit 610 is specifically used to input the multiple subject images into the encoding network to obtain multiple feature maps that correspond one-to-one with the multiple subject images; input the multiple feature maps into the classification network to obtain prediction information corresponding to each of the multiple subject images; and train the classification network based on the ground truth class labels corresponding to each of the multiple subject images, the prediction information corresponding to each of the multiple subject images, and the second loss function to obtain the target classification model. The second loss function is used to reduce the difference between the ground truth class labels corresponding to each of the multiple subject images and the prediction information corresponding to each of the multiple subject images.
[0103] As one approach, the training classification matting model further includes an upsampling network. The first annotation information includes multiple ground truth mask images. The model generation unit 610 is specifically used to input the multiple subject images into the encoding network to obtain multiple feature maps corresponding one-to-one with the multiple subject images; input the multiple feature maps into the classification network to obtain the predicted category corresponding to each of the multiple subject images; input the multiple feature maps into the decoding network to obtain multiple second prediction mask images corresponding one-to-one with the multiple subject images; and input the predicted category corresponding to each of the multiple subject images and the multiple second prediction mask images into the judgment module to obtain the target subject. The image is a subject image whose predicted category is determined by the judgment module to belong to the target category. The second predicted mask image corresponding to the subject image is input into the upsampling network to obtain a third predicted mask image, the resolution of which is higher than that of the second predicted mask image. The upsampling network is trained based on the third predicted mask image, the ground truth mask image corresponding to the third predicted mask image, and the third loss function to obtain the target classification matting model. The third loss function is used to reduce the pixel value difference between each pixel in the third predicted mask image and the corresponding ground truth mask image.
[0104] Optionally, the target category includes multiple target subcategories, and the upsampling network includes multiple upsampling subnetworks. Each upsampling subnetwork is used to output the output result corresponding to the target subcategory. The output result of each target subcategory has a different resolution. The model generation unit 610 is specifically used to input the predicted categories corresponding to the multiple subject images and the multiple second prediction mask images into the judgment module to obtain the target subject image. The target subject image is the subject image whose corresponding predicted category is determined by the judgment module to belong to any one of the multiple target subcategories. The second prediction mask image corresponding to the target subject image is input into the corresponding upsampling subnetwork to obtain the third prediction mask image.
[0105] Optionally, the model generation unit 610 is specifically used to acquire a target mask image, wherein the target mask image is an image with the same resolution as the third predicted mask image obtained based on the ground truth mask image corresponding to the third predicted mask image; the upsampling network is trained based on the third predicted mask image, the target mask image, and the third loss function to obtain the target classification matting model.
[0106] Please see Figure 8 This application provides an image processing apparatus 800, the apparatus 800 comprising:
[0107] The image acquisition unit 810 is used to acquire the image to be cut out.
[0108] The image matting result acquisition unit 820 is used to input the image to be matted into a target classification matting model to obtain the matting result of the image to be matted, wherein the target classification matting model is obtained based on any one of the methods described in claims 1-6.
[0109] In one approach, the matting result acquisition unit 820 is specifically used to input the image to be matted into the target classification matting model to obtain a mask image of the image to be matted; and to obtain the matting result based on the image to be matted and the mask image.
[0110] Optionally, the target classification matting model includes an encoding network, a decoding network, a classification network, a judgment module, and an upsampling network. The matting result acquisition unit 820 is specifically used to input the multiple subject images into the encoding network to obtain multiple feature maps corresponding one-to-one with the multiple subject images; input the multiple feature maps into the classification network to obtain the predicted category corresponding to each of the multiple subject images; input the multiple feature maps into the decoding network to obtain multiple reference prediction mask images corresponding one-to-one with the multiple subject images; input the predicted category corresponding to each of the multiple subject images and the multiple reference prediction mask images into the judgment module; if the judgment module determines that the predicted category is the target category, the reference prediction mask image is input into the upsampling network to obtain the mask image; if the judgment module determines that the predicted category is not the target category, the reference prediction mask image is used as the mask image.
[0111] The following will combine Figure 9 This application describes an electronic device.
[0112] Please see Figure 9 Based on the aforementioned model generation method, image processing method, and apparatus, this application embodiment also provides another electronic device 100 capable of executing the aforementioned model generation method and image processing method. The electronic device 100 includes one or more (only one shown) processors 102 and a memory 104 coupled together. The memory 104 stores programs capable of executing the contents of the aforementioned embodiments, and the processors 102 can execute the programs stored in the memory 104.
[0113] The processor 102 may include one or more processing cores. The processor 102 connects to various parts within the electronic device 100 using various interfaces and lines, and performs various functions and processes data by running or executing instructions, programs, code sets, or instruction sets stored in the memory 104, and by calling data stored in the memory 104. Optionally, the processor 102 may be implemented using at least one hardware form of Digital Signal Processing (DSP), Field-Programmable Gate Array (FPGA), or Programmable Logic Array (PLA). The processor 102 may integrate one or a combination of several of the following: Central Processing Unit (CPU), Graphics Processing Unit (GPU), and modem. The CPU primarily handles the operating system, user interface, and applications; the GPU is responsible for rendering and drawing the displayed content; and the modem handles wireless communication. It is understood that the modem may also not be integrated into the processor 102 and may be implemented separately using a communication chip.
[0114] The memory 104 may include random access memory (RAM) or read-only memory (ROM). The memory 104 can be used to store instructions, programs, code, code sets, or instruction sets. The memory 104 may include a program storage area and a data storage area. The program storage area may store instructions for implementing an operating system, instructions for implementing at least one function (such as touch functionality, sound playback functionality, image playback functionality, etc.), and instructions for implementing the various method embodiments described below. The data storage area may also store data created by the terminal 100 during use (such as phonebook data, audio and video data, chat log data, etc.).
[0115] Please refer to Figure 10 This diagram illustrates a structural block of a computer-readable storage medium provided in an embodiment of this application. The computer-readable storage medium 1000 stores program code that can be invoked by a processor to execute the methods described in the above method embodiments.
[0116] The computer-readable storage medium 1000 may be an electronic memory such as flash memory, EEPROM (Electrically Erasable Programmable Read-Only Memory), EPROM, hard disk, or ROM. Optionally, the computer-readable storage medium 1000 includes a non-transitory computer-readable storage medium. The computer-readable storage medium 1000 has storage space for program code 1010 that performs any of the method steps described above. This program code can be read from or written to one or more computer program products. The program code 1010 may be compressed, for example, in a suitable form.
[0117] In summary, the model generation method, image processing method, apparatus, and electronic device provided in this application, after training a matting model to be trained based on multiple subject images, a first training dataset containing first annotation information of each of the multiple subject images, and a first loss function to obtain a low-resolution matting model, then training a classification model to be trained based on multiple subject images identical to the first training dataset, a second training dataset containing second annotation information of each of the multiple subject images, and a second loss function to obtain a target classification model, and then training a classification matting model to be trained based on the first training dataset and a third loss function, including the encoding network and the encoding network of the low-resolution matting model, to obtain a target classification matting model, wherein the judgment module is used to determine the output result of the target classification matting model based on the classification result of the classification network, and the resolution of the output result corresponding to different classification results is different. The above method allows for the following steps: First, a low-resolution matting model is trained using a training dataset and a first loss function. Then, a classification model containing the low-resolution matting model is trained using the training dataset and a second loss function to obtain a target classification model. Finally, a classification matting model containing the low-resolution matting model, its encoding and decoding networks, and the target classification model's classification network is trained using the training dataset and a third loss function to obtain a target classification matting model. This allows the target classification matting model to output different resolutions for different types of images, thereby improving its matting flexibility and ensuring that its output meets the resolution requirements of different scenarios and tasks.
[0118] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application.
Claims
1. A model generation method, characterized in that, The method includes: The matting model to be trained is trained based on the first training dataset and the first loss function to obtain a low-resolution matting model. The first training dataset includes multiple subject images and the first annotation information of each of the multiple subject images. The target classification model is obtained by training the classification model to be trained based on the second training dataset and the second loss function. The second training dataset includes multiple subject images that are the same as the first training dataset, as well as the second annotation information of each of the multiple subject images. The classification model to be trained includes a classification network and the encoding network of the low-resolution matting model. The training dataset and the third loss function are used to train the classification matting model to obtain the target classification matting model. The classification matting model to be trained includes the encoding network, the decoding network of the low-resolution matting model, the classification network, and the judgment module. The judgment module is used to determine the output result of the target classification matting model based on the classification result of the classification network. The resolution of the output result corresponding to different classification results is different.
2. The method according to claim 1, characterized in that, The first annotation information includes multiple ground truth mask images. The step of training the matting model based on the first training dataset and the first loss function to obtain a low-resolution matting model includes: The resolution of the multiple ground truth mask images is reduced to obtain multiple low-resolution ground truth mask images; The multiple subject images are input into the matting model to be trained to obtain multiple first prediction mask images that correspond one-to-one with the multiple subject images; The matting model to be trained is obtained by training the multiple low-resolution ground truth mask images, the multiple first predicted mask images, and the first loss function, wherein the first loss function is used to reduce the pixel value difference between each pixel of the multiple low-resolution ground truth mask images and the multiple first predicted mask images.
3. The method according to claim 1, characterized in that, The second annotation information includes the true category labels corresponding to each of the multiple subject images. The step of training the classification model to be trained based on the second training dataset and the second loss function to obtain the target classification model includes: The multiple subject images are input into the coding network to obtain multiple feature maps that correspond one-to-one with the multiple subject images; The multiple feature maps are input into the classification network to obtain prediction information corresponding to each of the multiple subject images; The classification network is trained based on the true category labels corresponding to each of the multiple subject images, the predicted information corresponding to each of the multiple subject images, and the second loss function to obtain the target classification model. The second loss function is used to reduce the difference between the true category labels corresponding to each of the multiple subject images and the predicted information corresponding to each of the multiple subject images.
4. The method according to claim 1, characterized in that, The classification matting model to be trained further includes an upsampling network. The first annotation information includes multiple ground truth mask images. The step of training the classification matting model to be trained based on the first training dataset and the third loss function to obtain the target classification matting model includes: The multiple subject images are input into the coding network to obtain multiple feature maps that correspond one-to-one with the multiple subject images; The multiple feature maps are input into the classification network to obtain the predicted categories corresponding to each of the multiple subject images; The multiple feature maps are input into the decoding network to obtain multiple second prediction mask images that correspond one-to-one with the multiple main images; The predicted categories corresponding to each of the multiple subject images and the multiple second prediction mask images are input into the judgment module to obtain the target subject image. The target subject image is the subject image whose corresponding predicted category is determined by the judgment module to belong to the target category. The second prediction mask image corresponding to the target subject image is input into the upsampling network to obtain the third prediction mask image, wherein the resolution of the third prediction mask image is higher than that of the second prediction mask image; The upsampling network is trained based on the third predicted mask image, the ground truth mask image corresponding to the third predicted mask image, and the third loss function to obtain the target classification matting model. The third loss function is used to reduce the pixel value difference between each pixel in the third predicted mask image and the corresponding ground truth mask image.
5. The method according to claim 4, characterized in that, The target category includes multiple target subcategories, and the upsampling network includes multiple upsampling subnetworks. Each upsampling subnetwork is used to output the output result corresponding to the target subcategory. The resolution of the output result of each target subcategory is different. The step of inputting the predicted category corresponding to each of the multiple subject images and the multiple second prediction mask images into the judgment module to obtain the target subject image includes: The predicted categories corresponding to each of the multiple subject images and the multiple second prediction mask images are input into the judgment module to obtain the target subject image. The target subject image is the subject image whose corresponding predicted category is determined by the judgment module to belong to any one of the multiple target subcategories. The step of inputting the second prediction mask image corresponding to the target subject image into the upsampling network to obtain the third prediction mask image includes: The second prediction mask image corresponding to the target subject image is input into the corresponding upsampling sub-network to obtain the third prediction mask image.
6. The method according to claim 5, characterized in that, The upsampling network is trained based on the third predicted mask image, the ground truth mask image corresponding to the third predicted mask image, and the third loss function to obtain the target classification matting model, including: Obtain a target mask image, wherein the target mask image is an image with the same resolution as the third predicted mask image obtained based on the ground truth mask image corresponding to the third predicted mask image; The upsampling network is trained based on the third predicted mask image, the target mask image, and the third loss function to obtain the target classification matting model.
7. An image processing method, characterized in that, The method includes: Obtain the image to be cut out; The image to be cut out is input into a target classification cutout model to obtain the cutout result of the image to be cut out. The target classification cutout model is obtained based on any one of the methods described in claims 1-6.
8. The method according to claim 7, characterized in that, The step of inputting the image to be matted into a target classification matting model to obtain the matting result of the image to be matted includes: The image to be cut out is input into the target classification cutting out model to obtain the mask image of the image to be cut out; The cutout result is obtained based on the image to be cut out and the mask image.
9. The method according to claim 8, characterized in that, The target classification matting model includes an encoding network, a decoding network, a classification network, a judgment module, and an upsampling network. The step of inputting the image to be matted into the target classification matting model to obtain the output result of the image to be matted includes: The multiple subject images are input into the coding network to obtain multiple feature maps that correspond one-to-one with the multiple subject images; The multiple feature maps are input into the classification network to obtain the predicted categories corresponding to each of the multiple subject images; The multiple feature maps are input into the decoding network to obtain multiple reference prediction mask images that correspond one-to-one with the multiple main images; The predicted categories corresponding to each of the multiple subject images and the multiple reference prediction mask images are input into the judgment module. If the prediction category is determined to be the target category based on the judgment module, the reference prediction mask image is input into the upsampling network to obtain the mask image. If the judgment module determines that the predicted category is not the target category, the reference prediction mask image is used as the mask image.
10. A model generation apparatus, characterized in that, The device includes: The model generation unit is used to train a matting model to be trained based on a first training dataset and a first loss function to obtain a low-resolution matting model. The first training dataset includes multiple subject images and first annotation information for each of the multiple subject images. It also trains a classification model to be trained based on a second training dataset and a second loss function to obtain a target classification model. The second training dataset includes multiple subject images identical to those in the first training dataset and second annotation information for each of the multiple subject images. The classification model to be trained includes a classification network and an encoding network of the low-resolution matting model. Finally, it trains a classification matting model to be trained based on the first training dataset and a third loss function to obtain a target classification matting model. The classification matting model to be trained includes the encoding network, a decoding network of the low-resolution matting model, the classification network, and a judgment module. The judgment module is used to determine the output result of the target classification matting model based on the classification result of the classification network. Different classification results correspond to output results with different resolutions.
11. An image processing apparatus, characterized in that, The device includes: The image to be cut out acquisition unit is used to acquire the image to be cut out; The matting result acquisition unit is used to input the image to be matted into a target classification matting model to obtain the matting result of the image to be matted, wherein the target classification matting model is obtained based on any one of the methods described in claims 1-6.
12. An electronic device, characterized in that, Includes one or more processors and memory; One or more programs are stored in the memory and configured to be executed by the one or more processors, the one or more programs being configured to perform the method of any one of claims 1-9.
13. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores program code, wherein the method described in any one of claims 1-9 is executed when the program code is run.
Citation Information
Patent Citations
Image processing method and device and electronic equipment
CN114170250A
Loss value-based matting model training method and device, equipment and medium
CN114255378A