A method and device for automatically segmenting materials and refining local animation targets
By using the combination technology of the object detector Grounding DINO and a generative adversarial network in animation production, the accurate positioning and refinement of material targets in animation images is solved, and the problems of inaccurate segmentation and positioning and uncontrollable detail generation in the existing technology are solved, and the efficiency and quality of animation production are improved.
Patent Information
- Application Number
- CN202411903716.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-23
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2044-12-23
AI Technical Summary
The prior art cannot accurately segment and position the required material type in animation production, and the original design plan is easily changed when generating local details and the results are uncontrollable.
The material object detection model is constructed based on the object detector Grounding DINO, and the image refinement model is constructed in combination with the generative adversarial network. Through material object detection and image refinement processing, the material object detection and image refinement processing can be achieved accurately positioned and refined.
It improves the automatic positioning and segmentation accuracy of the material target area, reduces manual intervention, and the generated refined local material images are more in line with expectations, reducing the economic cost and time period of animation production.
Smart Images

Figure CN119359878B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of computer vision and image processing, and specifically relates to a method and device for automatically segmenting materials and refining local targets of animations. Background Art
[0002] The application of artificial intelligence in the field of animation production is becoming more and more popular. With the advancement of technology, artificial intelligence has begun to play an important role in various aspects of animation production. For example, artificial intelligence can help animators automate some repetitive tasks, such as color matching and image processing, which used to be done manually and time-consuming and labor-intensive.
[0003] AI can also be used in the design and layout of animation scenes. By using algorithms such as generative adversarial networks (GANs), a first draft of an animation scene can be automatically generated based on a given scene description. This can help animators save time and effort, and can create unique and interesting scene designs.
[0004] In the animation production process, image refinement is an important process in the production of two-dimensional plane animation. Because different methods are needed for fine depiction for different areas and targets, restoring the details of objects usually requires strong professional knowledge. Artificial painting will also fail to achieve the expected effect due to the lack of understanding of the creator's intention and the limitation of technical level, resulting in deviations. In addition, the depiction of details needs to take into account the specific material, color matching and texture details, resulting in a long time period and high economic cost for the production of two-dimensional plane animation. Therefore, research that can locate and depict high-quality details based on local position is both necessary and attractive.
[0005] With the rapid development of computer graphics and machine learning, the way of creating animation has also changed. Traditional manual drawing methods are being replaced by automated creation driven by deep learning. Deep learning models can generate high-quality animation works by learning and simulating the creative style of human artists. For example, the patent application with publication number CN118397135A discloses a method for consistent coloring of comic characters based on deep learning. For example, the patent application with publication number CN118470198A discloses an improved system for generating animation sketches based on deep learning.
[0006] The above existing methods mainly use automatic segmentation networks such as SAM (Segment Anything Model) to generate masks in the Stable Diffusion model, and then perform local characterization on the masked area by secondary editing. However, this method has two shortcomings. One is that it cannot accurately segment and locate according to the required categories. The other is that using a large model to generate local details will change the original design plan and the generated results are uncontrollable. Summary of the invention
[0007] In view of the above, the purpose of the present invention is to provide a method and device for automatically segmenting materials and refining local targets of animation, which can efficiently and accurately segment and extract local material targets from animation images under complex backgrounds, and further refine the local material targets to improve the detail display effect.
[0008] To achieve the above-mentioned purpose of the invention, an embodiment provides a method for automatically segmenting materials and refining local animation targets, comprising the following steps:
[0009] A material target detection model is built based on the target detector Grounding DINO, and the material target detection model is used to detect targets in animation images based on the input material classification text to obtain the material target box and the corresponding material category.
[0010] The target segmentation model is used to segment the material target of the animation image with the material target frame as the guiding information to obtain the mask map of the material target, and the animation local map corresponding to the material target is obtained from the animation image based on the mask map;
[0011] An image refinement model is constructed based on a generative adversarial network, and the image refinement model is used to refine the animation local map, and the generated refined local material image is fit to the animation image.
[0012] Preferably, a material object detection model is constructed based on the object detector Grounding DINO, including:
[0013] Prepare an animated image containing intrinsic colors, annotate a material target frame and a material category for a specified material in the animated image, and generate a content summary of the animated image to obtain sample data;
[0014] The parameters of the feature extraction module in Grounding DINO are fixed, and the animated image in the sample data is used as the image input, and the content summary is used as the text input. After the image features and text features are extracted by the feature extraction module, the features are enhanced and fused by the feature enhancement module. Then, the features are input to the language-guided query selection module to select features related to the input text as the decoding query, and the features of the decoding query are input to the cross-modal decoding module for regression prediction and material classification. Based on the predicted material target box and the corresponding material category, the logistic regression loss and classification loss are constructed to optimize the parameters of the feature enhancement module, the language-guided query selection module, and the cross-modal decoding module in Grounding DINO. The Grounding DINO with optimized parameters is used as the material target detection model.
[0015] Preferably, the target segmentation model includes an image encoder, a guidance information encoder, and a mask decoder. The image encoder encodes and embeds the input animation image to obtain an image embedding. The guidance information encoder is used to encode and embed the guidance information to obtain a guidance embedding. The mask decoder is used to decode and predict a mask image based on the image embedding and the guidance embedding.
[0016] Preferably, the image encoder adopts a ViT model pre-trained by masked auto-encoding (MAE);
[0017] In the guide information encoder, position encoding is performed on the point and frame information in the material target frame as the guide information, and the position encoding result is embedded as the guide;
[0018] The mask decoder includes a Transformer-based decoding module and a mask prediction head. The input image embedding and guide embedding are decoded by the decoding module and then predicted by the mask prediction head to obtain a mask image.
[0019] Preferably, an image refinement model is constructed based on a generative adversarial network, including:
[0020] Prepare an animated image containing intrinsic colors and its corresponding detailed image as sample data;
[0021] Construct a generative adversarial network consisting of a generator and a discriminator, where the generator is used to generate a refined local texture image after feature extraction based on the input animation image, and the discriminator is used to determine the authenticity of the input refined local texture image and the image with rich details;
[0022] Constructing a loss function including adversarial loss, perceptual loss between the generated refined local texture image and the detail-rich image as a label, style loss, and perceptual image patch similarity loss;
[0023] The loss function and sample data are used to optimize the parameters of the generative adversarial network, and the generator after parameter optimization is used as the image refinement model.
[0024] Preferably, in the generative adversarial network, the generator includes a first convolutional layer, multiple dense residual modules, an upsampling layer, a second convolutional layer, and a third convolutional layer, and a residual connection is provided between the output of the first convolutional layer and the input of the third convolutional layer. After the input animation image passes through the first convolutional layer and the dense residual module for feature extraction, the refined local texture image is output after upsampling and convolution calculation through the upsampling layer, the second convolutional layer and the third convolutional layer, and the discriminator adopts a U-Net structure.
[0025] Preferably, the adversarial loss includes a generator loss and a discriminator loss;
[0026] For perceptual loss, the feature maps of each layer extracted before the activation layer of the VGG network are used, and the weighted sum of the differences between all layer feature maps of the two images is used as the perceptual loss.
[0027] For style loss, we use the feature maps of each layer extracted before the activation layer of the VGG network from the refined local texture image and the detail-rich image, as well as the Gram matrix of each layer of the feature map, and take the sum of the differences between all the Gram matrices of the two images as the style loss.
[0028] Aiming at the perceptual image block similarity loss, the Alexnet structure is used to learn and refine the similarity scores between local texture images and detail-rich images as the perceptual image block similarity loss.
[0029] To achieve the above-mentioned purpose of the invention, an embodiment of the present invention further provides a device for automatically segmenting materials and refining local targets of animations, comprising:
[0030] The material target determination module is used to build a material target detection model based on the target detector Grounding DINO, and use the material target detection model to perform target detection on the animation image based on the input material classification text to obtain the material target frame and the corresponding material category;
[0031] A material target segmentation module is used to segment the material target of the animation image using the target segmentation model and the material target frame as the guide information to obtain a mask map of the material target, and obtain a partial animation map corresponding to the material target from the animation image based on the mask map;
[0032] The material target refinement module is used to build an image refinement model based on a generative adversarial network, use the image refinement model to refine the animation local image, and fit the generated refined local material image to the animation image.
[0033] To achieve the above-mentioned purpose of the invention, an embodiment further provides a computing device, including a memory and one or more processors, wherein the memory stores executable code, and when the one or more processors execute the executable code, they are used to implement the above-mentioned method of automatically segmenting materials and refining local animation targets.
[0034] To achieve the above-mentioned purpose of the invention, an embodiment further provides a computer-readable storage medium on which a program is stored. When the program is executed by a processor, the above-mentioned method of automatically segmenting materials and refining local animation targets is implemented.
[0035] Compared with the prior art, the present invention has the following beneficial effects:
[0036] The present invention constructs a target detection model by transferring the learning target detector Grounding DINO, so that the target detection model can accurately obtain the corresponding material target area through the unrefined animation image with only intrinsic color and the material classification text as the prompt word, further improving the accuracy of automatic positioning and segmentation of the material target area, and reducing the need for manual intervention;
[0037] The present invention trains an image refinement model for refining a specified material based on a generative adversarial network. The image refinement model can learn refined image features. Users can optimize the texture details of the animated image in a personalized way according to the material conditions they actually use, and can obtain a refined local material image that is more in line with expectations.
[0038] The present invention can complete the segmentation and refinement process of the local image corresponding to the target material only by using the input material classification text as a prompt, thereby reducing the production cost and cycle of manual painting, assisting the animation design process, and improving the work efficiency of animation drawing. BRIEF DESCRIPTION OF THE DRAWINGS
[0039] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.
[0040] Figure 1 is a flow chart of a method for automatically segmenting materials and refining local animation targets provided by an embodiment;
[0041] Figure 2 It is a flowchart of building a material target detection model based on the target detector Grounding DINO provided in the embodiment;
[0042] Figure 3 It is a schematic diagram of the structure and process of the image refinement model provided in the embodiment;
[0043] Figure 4 It is a schematic diagram of the structure of the device for automatically segmenting materials and refining local targets of animations provided in an embodiment. DETAILED DESCRIPTION
[0044] To make the purpose, technical solution and advantages of the present invention more clearly understood, the present invention is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific implementation methods described herein are only used to explain the present invention and do not limit the scope of protection of the present invention.
[0045] The inventive concept of the present invention is: in view of the technical problem that the prior art cannot accurately segment and locate according to the required material type, the present invention constructs a material target detection model based on the target detector Grounding DINO, and can accurately locate and segment the target material type only according to the input material classification text as a prompt word. At the same time, in view of the technical problem that the prior art will change the original design scheme and the result is uncontrollable when generating local details, the present invention obtains the animation local map corresponding to the target material by positioning and segmenting, and optimizes the generative adversarial network in a personalized and refined manner by using the specified material image, so that the obtained image refinement model can optimize the texture details of the animation image in a personalized manner according to the actual material situation, thereby obtaining a more expected refined local material image.
[0046] like Figure 1 As shown, based on the above inventive concept, an embodiment provides a method for automatically segmenting materials and refining local animation targets, comprising the following steps:
[0047] S1, build a material target detection model based on the target detector Grounding DINO, and use the material target detection model to perform target detection on the animation image based on the input material classification text to obtain the material target box and the corresponding material category.
[0048] In the embodiment, the target detector Grounding DINO is selected as the basic model, and the target detector Grounding DINO is transferred and learned through sample data for material target detection to build a material target detection model. The specific process is as follows:
[0049] Sample data construction: Prepare and organize animation data sets and prepare animation images with intrinsic colors. The size of the animation images can be 1024×968 pixels. Use Roboflow annotation tools to annotate material target boxes and material categories for specified materials in the animation images. The material target boxes are annotated in the form of coordinates. At the same time, a content summary of the animation image is generated to obtain sample data and save it in ODVG (Object Detection Visual Grounding) data format.
[0050] like Figure 2 As shown in the figure, Grounding DINO includes a feature extraction module, a feature enhancement module, a language-guided query selection module, and a cross-modal decoding module. The feature extraction module includes a text feature extraction submodule and an image feature extraction submodule. When training Grounding DINO, first, the feature extraction module parameters in Grounding DINO are fixed, because they have been pre-trained on a larger dataset and can be used to represent image and text features well. Therefore, the text feature extraction submodule and the image feature extraction submodule are frozen during the training process.
[0051] Then, the animated image in the sample data is used as the image input, and the content summary is used as the text input. After the image features and text features are extracted by the feature extraction module respectively, the feature enhancement module is used to perform feature enhancement and fusion of the two modal features of text features and image features. The feature fusion adopts cross-attention from image to text and text to image to perform feature fusion.
[0052] Next, the fused features are input into the language-guided query selection module to select features related to the input text as cross-modal decoding queries to achieve modality alignment, and the features of the decoding queries are input into the cross-modal decoding module for target regression prediction and material classification to obtain the predicted material target box and the corresponding material category.
[0053] Finally, based on the predicted material target box and the corresponding material category, logistic regression loss and classification loss are constructed for transfer learning to optimize the feature enhancement module, language-guided query selection module, and cross-modal decoding module parameters in Grounding DINO. The parameter-optimized Grounding DINO is used as a material target detection model.
[0054] For the logistic regression loss about the bounding box, L1 loss and GIOU loss are used. Focal Loss is used to match the material category and the predicted target. In detail, the query material text and the result of the model feature output layer are dot-multiplied; then, the predicted box and the predicted material classification are paired with the real box and the real category through binary matching; so the final classification loss is calculated based on the loss between the matched result and the real value. In order to make the model results more convincing, auxiliary losses are added to each decoding layer for constraint.
[0055] During training, the learning rate is a hyperparameter of the optimizer that controls the pace of model updates. Since small-scale data transfer learning is performed, the learning rate is set to 0.0001 to prevent unstable model training. The model learning rate scheduler uses the OneCycleLR method. In OneCycleLR, the learning rate first increases linearly to the maximum value, and then gradually decreases through the decay and cosine annealing stages, reducing the risk of overfitting and improving the efficiency of model training. The optimizer uses AdamW.
[0056] The embodiment also provides the comparison results of material classification before and after transfer learning of the object detector Grounding DINO, as shown in Table 1:
[0057] Table 1
[0058] ;
[0059] It can be seen from Table 1 that the accuracy of image classification of target material categories after transfer learning is much higher than before transfer learning, which is sufficient to prove that transfer learning of the target detector Grounding DINO can improve the classification accuracy of target data.
[0060] S2, using the target segmentation model to use the material target frame as the guiding information, to perform material target segmentation on the animation image, to obtain the mask map of the material target, and to obtain the animation local map corresponding to the material target from the animation image based on the mask map.
[0061] In the embodiment, the material target frame determined by target detection in S1 is used as guidance information for prompting segmentation, and the input animation image is segmented into material targets based on the guidance information to obtain a mask image of the material target corresponding to the material target frame.
[0062] In an embodiment, the target segmentation model includes an image encoder, a guide information encoder, and a mask decoder. The image encoder encodes and embeds the input animation image to obtain an image embedding. Specifically, a ViT (Vision Transformer) model pre-trained by masked auto-encoder MAE (masked auto-encoder) can be used. After fine-tuning the ViT model, high-resolution input images can be processed. The number of parameters used to embed the input image is relatively large, but only one vector embedding needs to be calculated for the same input image. The image embedding can be reused for different prompt information, thereby reducing the inference pressure. The guide information encoder is used to encode and embed the guide information to obtain a guide embedding. Specifically, the point and box information in the material target box of the guide information can be used for position encoding, and the position encoding result is used as the guide embedding. The mask decoder is used to decode and predict the mask map based on the image embedding and the guide embedding. The specific mask decoder includes a Transformer-based decoding module and a mask prediction head. The input image embedding and guide embedding are decoded by the decoding module and then predicted by the mask prediction head.
[0063] After obtaining the mask image, an animation local image corresponding to the material target is obtained from the animation image based on the mask image. The animation local image is the local original image segmented from the animation image, which serves as the basis for subsequent image refinement.
[0064] S3, builds an image refinement model based on a generative adversarial network, uses the image refinement model to refine the animation local map, and fits the generated refined local material image to the animation image.
[0065] In an embodiment, an image refinement model is constructed based on a generative adversarial network, and the number of animation local images separated from the original animation image based on the mask image is input into the image refinement model. A refined local material image is generated after calculation, and the generated refined local material image replaces the material at the original inherent color position to obtain a more refined material picture.
[0066] The trained generative adversarial network can generate a refined image with rich details and colors based only on the input intrinsic color image. The training process of the generative adversarial network is:
[0067] First, prepare animated images with intrinsic colors and their corresponding detailed images as sample data, where the animated images with intrinsic colors are used as input images and the detailed images are used as label images. Specifically, about 850 animated images with intrinsic colors of different materials such as wood, stone, brick, plant, wall, metal, etc. are prepared in RGB format. These animated images are cut into small blocks of 128×128 size to adapt to the input size of the generative adversarial network. At the same time, these image blocks do not need any preprocessing to prevent the loss of additional information. The detailed images obtained after the intrinsic color animation images are portrayed have a large variation of detail richness, so as to increase the generalization ability of the trained model.
[0068] Then, a generative adversarial network consisting of a generator and a discriminator is constructed, wherein the generator is used to generate a refined local texture image after feature extraction based on the input animation image, and the discriminator is used to distinguish the authenticity of the input refined local texture image and the image with rich details. Figure 3 As shown in the figure, the generator uses dense residual components such as Densnet as the basic components of the generator. Dense connections are all performed by channel fusion, which mainly means that the current output requires the output of all previous layers. At the same time, in order to stabilize the training and make the generator have better generalization, the use of BN (BatchNomalization) is removed. Based on this, the generator includes the first convolutional layer, multiple dense residual modules, an upsampling layer, a second convolutional layer, and a third convolutional layer, and a residual connection is provided between the output of the first convolutional layer and the input of the third convolutional layer. After the input animation image is extracted by the first convolutional layer and the dense residual module, it is output based on the upsampling and convolution calculations of the upsampling layer, the second convolutional layer, and the third convolutional layer to refine the local texture image.
[0069] In the generator, the image folding method is used to reduce the width and height of the image and increase the number of channels of the image to reduce the complexity of the calculation. Multiple dense residual modules are used for feature extraction to enhance the image's ability to capture details. The specific residual dense module is composed of multiple residual densely connected convolutional layers, which are used to capture deeper features and improve the performance of the network. After passing through a series of residual dense modules, the network will enlarge the size of the feature map to the size of the target high-resolution image through the upsampling layer, and finally convert the upsampled feature map into a high-resolution refined local material image through the convolution layer.
[0070] The discriminator adopts the U-Net structure. Because of its powerful spatial capture capability, it can more accurately evaluate images and local texture details. Accurate gradient feedback can also help the generator learn how to reduce artifacts and produce natural high-resolution images.
[0071] Next, the parameters of the constructed generative adversarial network are initialized, and a loss function is constructed, which includes adversarial loss, perceptual loss based on the generated refined local material image and the detail-rich image as a label, style loss, and perceptual image block similarity loss.
[0072] For perceived loss , the feature maps of each layer extracted before the activation layer of the VGG network are used to refine the local texture image and the detail-rich image, and the weighted sum of the differences between the feature maps of all layers of the two images is used as the perceptual loss; the feature map before the activation layer of the VGG network is used because the feature map before activation has more detailed details and can bring stronger supervision. Suppose the following feature layers conv1, conv2, conv3, conv4, conv5 of the VGG19 network are selected, where the weight of each feature layer is The values are 0.1, 0.1, 0.25, 0.5, 0.5, respectively, then the perceptual loss It is expressed as:
[0073] ;
[0074] in, represents the feature map index, N represents the total number of feature maps used, and in the embodiment, N is set to 5. Represents the refined local material image output by the generator At the feature level The feature map obtained in Represents detail-rich images as labels At the feature level The feature map obtained in represents the L2 norm;
[0075] For style loss , the feature maps of each layer extracted before the activation layer in the VGG network are extracted from the refined local texture image and the detail-rich image, as well as the Gram matrix of each layer of the feature map. The Gram matrix is the covariance matrix of the feature layer vector, and the sum of the differences between all the layer Gram matrices of the two images is used as the style loss. , expressed as:
[0076] ;
[0077] in, Represents the refined local material image output by the generator At the feature level The Gram matrix of the feature map obtained in Represents detail-rich images as labels At the feature level The Gram matrix of the feature map obtained in;
[0078] For perceptual image patch similarity loss , the Alexnet structure is used to learn the similarity score between the local texture image and the detail-rich image as the perceptual image block similarity loss. The lower the similarity score, the more similar the two images are. express:
[0079] ;
[0080] in, Represents the refined local material image output by the generator In the Alexnet structure i Activation Layer The activation value obtained in Represents detail-rich images as labels In the activation layer of the Alexnet structure The activation value obtained in represents the scaling parameters learned from the data to adjust the impact of each layer of features, represents the weight of each activation layer.
[0081] The purpose of adversarial loss is to make the judgment distribution of the real detail-rich image minus the average distribution of the generated refined local texture image, and then perform sigmoid processing on the above result to make the result closer to 1; let the judgment distribution of the generated refined local texture image minus the average distribution of the detail-rich image, and then perform sigmoid processing on the above result to make the result closer to 0. Adversarial loss consists of two parts: generator loss and the discriminator loss .
[0082] The goal of the generator is to generate fake images such that the discriminator cannot distinguish between fake images and real images. Therefore, the generator loss Defined as:
[0083] ;
[0084] in, Represents a partial animation The distribution of , represents the refined local material image generated by the generator, is the discriminator’s judgment result on the refined local material image. Expressing hope;
[0085] The goal of the discriminator is to distinguish between real detail-rich images and generated refined local texture images. Therefore, the discriminator loss It can be defined as:
[0086] ;
[0087] in, Represents detailed images The distribution of ;
[0088] The total loss function is for:
[0089] ;
[0090] in, Indicates the weight, the value can be 0.1, Represents the weight, which can be 5.0. The smaller weight of the perceptual loss here helps to improve the stability of the model and avoid focusing too much on the content in the early stage of training and affecting the final effect. The weight of the style loss is chosen to be 5.0 so that the model can better match the style index and pursue better artistic effects.
[0091] Finally, the loss function and sample data are used to optimize the parameters of the generative adversarial network. The optimal network weights are obtained by training to continuously approach the minimization result. The generator with optimized parameters is used as the image refinement model.
[0092] When using the image refinement model for image refinement applications, the animation local image separated from the original animation image according to the mask image is input into the image refinement model to perform refinement processing on the specific material, and a refined local material image is obtained. Then the original animation image is read, and the refined local material image is pasted back to the original image for display processing. Specifically, the original inherent color animation image, the mask image, and the generated refined local material image are read, and the OpenCV image processing function is used. The mask area is determined by applying a threshold, and the noise is removed by corrosion and expansion. The mask area of the generated refined local material image is extracted, and a new image of the same size as the original image is created. The extracted refined local material image is pasted onto the new image and the pasted image is saved.
[0093] The method of the present invention obtains the generation of corresponding material areas and mask images by means of transfer learning, locates the local area that needs to be refined, generates and returns the locally refined target image through a refinement network.
[0094] like Figure 4As shown, the embodiment also provides a device 40 for automatically segmenting materials and refining local targets of animations, including a material target determination module 41, a material target segmentation module 42, and a material target refinement module 43, wherein the material target determination module 41 is used to build a material target detection model based on the target detector Grounding DINO, and use the material target detection model to perform target detection on the animation image based on the input material classification text to obtain a material target frame and a corresponding material category; the material target segmentation module 42 is used to use the target segmentation model to use the material target frame as guiding information to perform material target segmentation on the animation image to obtain a mask map of the material target, and obtain an animation local map corresponding to the material target from the animation image based on the mask map; the material target refinement module 43 is used to build an image refinement model based on a generative adversarial network, use the image refinement model to refine the animation local map, and fit the generated refined local material image to the animation image.
[0095] It should be noted that the device for automatically segmenting materials and refining local targets of animations provided in the above-mentioned embodiments should be illustrated by the division of the above-mentioned functional modules when performing automatic material segmentation and local target refinement. The above-mentioned functions can be assigned to different functional modules as needed, that is, the internal structure of the terminal or server is divided into different functional modules to complete all or part of the functions described above. In addition, the device for automatically segmenting materials and refining local targets of animations provided in the above-mentioned embodiments and the method embodiment for automatically segmenting materials and refining local targets of animations belong to the same concept. The specific implementation process is detailed in the method embodiment for automatically segmenting materials and refining local targets of animations, which will not be repeated here.
[0096] Based on the same inventive concept, an embodiment further provides a computing device, including a memory and one or more processors, wherein executable code is stored in the memory, and when the one or more processors execute the executable code, the method for automatically segmenting materials and refining local animation targets is implemented, specifically including the following steps:
[0097] S1, build a material target detection model based on the target detector Grounding DINO, and use the material target detection model to perform target detection on the animation image based on the input material classification text to obtain the material target box and the corresponding material category;
[0098] S2, using the target segmentation model to use the material target frame as the guide information to segment the material target of the animation image, obtain the mask map of the material target, and obtain the animation local map corresponding to the material target from the animation image based on the mask map;
[0099] S3, builds an image refinement model based on a generative adversarial network, uses the image refinement model to refine the animation local map, and fits the generated refined local material image to the animation image.
[0100] The computing device provided in the embodiment, in addition to the processor and memory, also includes hardware required for other services such as internal bus, network interface, memory, etc. at the hardware level. The memory is a non-volatile memory, and the processor reads the corresponding computer program from the non-volatile memory into the memory and then runs it to implement the method of automatically segmenting materials and refining local animation targets described in S1-S3 above. Of course, in addition to the software implementation method, the present invention does not exclude other implementation methods, such as logic devices or a combination of software and hardware, etc., that is to say, the execution subject of the following processing flow is not limited to each logic unit, but can also be hardware or logic devices.
[0101] Based on the same inventive concept, an embodiment further provides a computer-readable storage medium on which a program is stored. When the program is executed by a processor, the method for automatically segmenting materials and refining local animation targets is implemented, specifically including the following steps:
[0102] S1, build a material target detection model based on the target detector Grounding DINO, and use the material target detection model to perform target detection on the animation image based on the input material classification text to obtain the material target box and the corresponding material category;
[0103] S2, using the target segmentation model to use the material target frame as the guide information to segment the material target of the animation image, obtain the mask map of the material target, and obtain the animation local map corresponding to the material target from the animation image based on the mask map;
[0104] S3, builds an image refinement model based on a generative adversarial network, uses the image refinement model to refine the animation local map, and fits the generated refined local material image to the animation image.
[0105] In the embodiment, computer-readable media include permanent and non-permanent, removable and non-removable media, and information storage can be implemented by any method or technology. The information can be computer-readable instructions, data structures, program modules or other data.
[0106] The specific implementation methods described above provide a detailed description of the technical solutions and beneficial effects of the present invention. It should be understood that the above is only the most preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, supplements and equivalent substitutions made within the scope of the principles of the present invention should be included in the protection scope of the present invention.
Claims
1. A method for automatically segmenting materials and refining local animation targets, characterized in that: The following steps are involved: A material target detection model is built based on the target detector Grounding DINO, and the material target detection model is used to detect targets in animation images based on the input material classification text to obtain the material target box and the corresponding material category. The target segmentation model is used to segment the material target of the animation image with the material target frame as the guiding information to obtain the mask map of the material target, and the animation local map corresponding to the material target is obtained from the animation image based on the mask map; An image refinement model is constructed based on a generative adversarial network. The image refinement model is used to refine the animation local image, and the generated refined local material image is fitted to the animation image. Among them, a material target detection model is constructed based on the target detector Grounding DINO, including: preparing an animated image with inherent color, and annotating the material target box and material category for the specified material in the animated image, and generating a content summary of the animated image to obtain sample data; fixing the parameters of the feature extraction module in Grounding DINO, taking the animated image in the sample data as the image input, and taking the content summary as the text input, respectively extracting image features and text features through the feature extraction module, and then performing feature enhancement and fusion through the feature enhancement module, and then inputting them into the language-guided query selection module to select features related to the input text as decoding queries, and inputting the features of the decoding queries into the cross-modal decoding module for regression prediction and material classification, and constructing a logistic regression loss and a classification loss based on the predicted material target box and the corresponding material category to optimize the parameters of the feature enhancement module, the language-guided query selection module, and the cross-modal decoding module in Grounding DINO, and the Grounding DINO after parameter optimization is used as a material target detection model; Among them, an image refinement model is constructed based on a generative adversarial network, including: preparing an animated image containing inherent colors and its corresponding detail-rich image as sample data; constructing a generative adversarial network including a generator and a discriminator, wherein the generator is used to generate a refined local material image after feature extraction based on the input animated image, and the discriminator is used to distinguish the authenticity of the input refined local material image and the detail-rich image; constructing a loss function, which includes an adversarial loss, a perceptual loss based on the generated refined local material image and the detail-rich image as a label, a style loss, and a perceptual image block similarity loss; using the loss function and the sample data to optimize the parameters of the generative adversarial network, and the generator after parameter optimization is used as the image refinement model.
2. The method for automatically segmenting materials and refining local animation targets according to claim 1, characterized in that: The target segmentation model includes an image encoder, a guide information encoder, and a mask decoder. The image encoder encodes and embeds the input animation image to obtain an image embedding. The guide information encoder is used to encode and embed the guide information to obtain a guide embedding. The mask decoder is used to decode and predict a mask image based on the image embedding and the guide embedding.
3. The method for automatically segmenting materials and refining local animation targets according to claim 2, characterized in that: The image encoder adopts the ViT model pre-trained by masked automatic encoding (MAE); In the guide information encoder, position encoding is performed on the point and frame information in the material target frame as the guide information, and the position encoding result is embedded as the guide; The mask decoder includes a Transformer-based decoding module and a mask prediction head. The input image embedding and guide embedding are decoded by the decoding module and then predicted by the mask prediction head to obtain a mask image.
4. The method for automatically segmenting materials and refining local animation targets according to claim 1, characterized in that: In the generative adversarial network, the generator includes a first convolutional layer, multiple dense residual modules, an upsampling layer, a second convolutional layer, and a third convolutional layer, and a residual connection is provided between the output of the first convolutional layer and the input of the third convolutional layer. After the input animation image is subjected to feature extraction by the first convolutional layer and the dense residual module, a refined local texture image is output after upsampling and convolution calculation by the upsampling layer, the second convolutional layer, and the third convolutional layer. The discriminator adopts a U-Net structure.
5. The method for automatically segmenting materials and refining local animation targets according to claim 1, characterized in that: The adversarial loss includes a generator loss and a discriminator loss; For perceptual loss, the feature maps of each layer extracted before the activation layer of the VGG network are used, and the weighted sum of the differences between all layer feature maps of the two images is used as the perceptual loss. For style loss, we use the feature maps of each layer extracted before the activation layer of the VGG network from the refined local texture image and the detail-rich image, as well as the Gram matrix of each layer of the feature map, and take the sum of the differences between all the Gram matrices of the two images as the style loss. Aiming at the perceptual image block similarity loss, the Alexnet structure is used to learn and refine the similarity scores between local texture images and detail-rich images as the perceptual image block similarity loss.
6. A device for automatically segmenting materials and refining local targets of animations, characterized in that: include: The material target determination module is used to build a material target detection model based on the target detector Grounding DINO, and use the material target detection model to perform target detection on the animation image based on the input material classification text to obtain the material target frame and the corresponding material category; Among them, a material target detection model is constructed based on the target detector Grounding DINO, including: preparing an animated image with inherent color, and annotating the material target box and material category for the specified material in the animated image, and generating a content summary of the animated image to obtain sample data; fixing the parameters of the feature extraction module in Grounding DINO, taking the animated image in the sample data as the image input, and taking the content summary as the text input, respectively extracting image features and text features through the feature extraction module, and then performing feature enhancement and fusion through the feature enhancement module, and then inputting them into the language-guided query selection module to select features related to the input text as decoding queries, and inputting the features of the decoding queries into the cross-modal decoding module for regression prediction and material classification, and constructing a logistic regression loss and a classification loss based on the predicted material target box and the corresponding material category to optimize the parameters of the feature enhancement module, the language-guided query selection module, and the cross-modal decoding module in Grounding DINO, and the Grounding DINO after parameter optimization is used as a material target detection model; A material target segmentation module is used to segment the material target of the animation image using the target segmentation model and the material target frame as the guide information to obtain a mask map of the material target, and obtain a partial animation map corresponding to the material target from the animation image based on the mask map; A material target refinement module is used to build an image refinement model based on a generative adversarial network, use the image refinement model to refine the animation local image, and fit the generated refined local material image to the animation image; Among them, an image refinement model is constructed based on a generative adversarial network, including: preparing an animated image containing inherent colors and its corresponding detail-rich image as sample data; constructing a generative adversarial network including a generator and a discriminator, wherein the generator is used to generate a refined local material image after feature extraction based on the input animated image, and the discriminator is used to distinguish the authenticity of the input refined local material image and the detail-rich image; constructing a loss function, which includes an adversarial loss, a perceptual loss based on the generated refined local material image and the detail-rich image as a label, a style loss, and a perceptual image block similarity loss; using the loss function and the sample data to optimize the parameters of the generative adversarial network, and the generator after parameter optimization is used as the image refinement model.
7. A computing device comprising a memory and one or more processors, wherein the memory stores executable code, characterized in that: When the one or more processors execute the executable code, they are used to implement the method for automatically segmenting materials and refining local animation targets according to any one of claims 1 to 5.
8. A computer-readable storage medium, characterized in that: A program is stored thereon, and when the program is executed by a processor, the method of automatically segmenting materials and refining local animation targets as described in any one of claims 1 to 5 is implemented.
Citation Information
Patent Citations
Cartoon character consistency coloring method and system based on deep learning
CN118397135A
Improved system for cartoon sketch generation based on deep learning
CN118470198A
Method for converting three-dimensional scene from rendering reality to physical reality
CN114842116A
Infrared small target detection method based on scene text information guidance
CN118762364A